Written from the repo, not summary -- every claim grep- or run-verified,
UNKNOWN where the repo cannot say.
State: HEAD and gitea main in sync at 71d3b7b, tree clean, 4,539 tests
passing, web build exit 0. Push remote is gitea, never origin.
DONE this arc, existence-verified: E1 movement strip, F9-F11 offseason hub
shell, E10 Report template, E12 /report archive, content engine + studio
API + preview, Wave-D1 motion primitives, honest served grade, archetype
doctrine. Also verified done and previously mis-boarded: D1 glyphs (39 of
83 -- every one that maps to a real archetype), A1 card token, B1 boundary
channel, book comparison, THE WIRE.
GATED with each gate named and each absence verified by 0-file grep: F5
article media and E16/F8 on the card-system reconciliation, in-season hub
IA on the content formula, E9/E15 on model accrual, E2/E6 on licensing,
E3 crown on measurement, E13 on another order. NO UNGATED WAVE-2 TARGETS
REMAIN -- the next move is a decision, not a build.
The two open decisions are stated with what each unblocks. Card-system
reconciliation is the cheapest: it frees F5 and E16/F8 together, and it
exists because I built a 1080x1350 renderer without checking whether a
designed card system existed. It did.
Accrual clock recorded at 0 eligible dates on all four items, with the
first trigger (10 calibration dates) and the attempt-floor-not-trust-floor
caveat carried forward.
Standing doctrine carried: Truth Law including designer samples and the
inverse Number(null) breach (zero is a real fact); the 83-glyph taxonomy
with 44 dormant slots; the repaired champion and the standing flag that
every prior factor verdict was measured against the broken baseline;
calibration withdrawn; unissuable A with a B+ 0.663 ceiling; both CI
guards; and compose-don't-fork, which this arc failed three times and
caught twice.
Both parallel workstreams written as seed prompts.
specs/HANDOFF-2026-08-08.md
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — the spec, read not recalled. E10: "Hybrid: dark billboard header
that survives every client, light paper body Gmail can't wreck. 600px,
stacked, no webfont dependence." Content law: "One email per slate day.
Top read, what changed, the record. Nothing else." E12: "EVERY ISSUE SHOWS
ITS OWN DAY RECORD -- THE ARCHIVE IS A LEDGER TOO."
COMPOSED, NOT FORKED. The audit had E10 as PARTIAL, not absent:
newsletterService already builds the daily report's CONTENT and lints its
voice. What was missing is the designed hybrid SHELL, so reportTemplate.js
is a template over that builder rather than a second report -- the same
call made for the movement strip, and for the same reason.
PHASE 1 — the hybrid shell is an ENGINEERING constraint, not a look, and
the tests say so: Gmail strips style blocks, Outlook ignores flexbox, and
a dark body renders as a black rectangle in several clients. Hence tables,
inline styles, 600px fixed, system fonts, no image required to read, and
the green SHIFTS from #00D4A0 to #00A57D on paper because the dark-mode
green is unreadable there.
FACT-CONTRACTED: a section whose data is absent is OMITTED and NAMED in
`omitted`, never filled. There is no code path producing a placeholder
figure. The honesty block carries the real numbers -- graded count,
cleared-ceiling count, the realized rate against baseline, and that we do
not issue A grades.
E1'S LAW TRAVELS EVEN THOUGH ITS RENDERING CANNOT. An SVG strip is not
reliable in email, so movementText carries the RULE: green only when the
move favours the read, and a flat market says FLAT · [N]D rather than
showing nothing.
NO DESIGNER SAMPLE DATA. Nabers 1,120.5, No 128, DAY RECORD 9-4 are a spec
for what a live issue renders; pasting them in would be fabrication
carrying a designer's authority and would look entirely correct. Tested.
PHASE 2 — /report is now the real archive, REPLACING the S41 redirect to
/blog. That redirect existed because the surface did not; E12 built it, so
the placeholder is correctly gone and the S41 test is updated rather than
worked around. Every row carries its own day record, and an unknown record
says UNSETTLED -- never a dash that reads as zero. Empty archive is an
honest state.
Backend: public read-only /api/report over Redis issues, plus the Next
proxy. Both surfaces registered under the reachability guard.
A test bug I made twice now: my check for forbidden sample values matched
the template's own doc block, which NAMES those values as things never to
paste. Documentation worth keeping, so both suites strip comments before
matching -- a guard that reads its own warning is not reading the code.
WAVE-2 STATUS: E1, F9-F11, E10, E12 done. Still gated -- F5 article media
and E16/F8 on the card-system reconciliation; the in-season hub IA on the
social chat's formula; E9/E15 on model; E2/E6 on licensing.
Read-only throughout; serving fingerprint unchanged including
newsletterService; accrual clock unchanged at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — specs/ARCHETYPE-TAXONOMY-DOCTRINE.md records the ruling as
shared law: 83 designed glyphs are the full four-sport taxonomy; a glyph
renders ONLY where its archetype is modeled and proven. 39 of 83 map to a
real archetype and are wired; the 44 unmapped are DORMANT SLOTS for
WNBA/NBA/Soccer, not a wiring gap. Wiring them would mean inventing 44
archetypes to consume artwork -- decoration presented as classification,
which is forbidden. DUAL THREAT and PAINT BOSS are modeled archetypes with
no mark: the mirror gap, flagged to the design side. When a sport's
archetypes ship, activation is a MANIFEST lookup, not new art.
PHASE 1 — E1 movement strip. The spec's own line is "the movement strip is
defined once here and reused everywhere a line has a past", so it is a
primitive, not a fourth chart.
RECONCILED RATHER THAN FORKED: lib/gradeShift.js ALREADY implements E1's
colour law -- toward/against/flat, including the direction flip that makes
an UNDER's favourable move the opposite sign of an OVER's. MovementStrip
CONSUMES buildGradeTimeline instead of reimplementing it, and a test
asserts it never redefines isUnder. GradeShift stays the grade-history
view; this is the reusable strip. That is the card-fork lesson applied
before it could happen again.
Spec laws honoured: STEPS NOT CURVES (H then V, no smoothing -- a curve
invents prices that never traded, and a test rejects any C/S/Q/T command);
green only when the move FAVOURS the read; FLAT renders as a hairline plus
FLAT · [N]D because a flat market is a finding; and too little history
says NO MOVEMENT HISTORY rather than rendering blank.
PHASE 2 — F9-F11 offseason hub shell, built from Vyndr Offseason.dc.html.
The spec's load-bearing words are used verbatim: "OUTLOOKS REPRICE ON NEWS
· NOT GAME ODDS" (an offseason number is not a game line), the QUIET WIRE
empty state ("No outlook-moving news since X. We don't manufacture
movement."), WHAT CHANGED TODAY as the hero with the countdown ambient and
top-right, the tag-colour-is-meaning row anatomy, the open -> NOW -> FAIR
triplet with the movement strip embedded, and the OUTLOOK ONLY block where
every row carries NOT GRADED.
THE DESIGN FILE'S SAMPLE DATA IS NOT IN THE COMPONENT. Wembanyama +420 ->
+330, Nabers cleared 11:42 AM, the Summer League names -- all of it is a
SPEC for what a live feed renders, and copying it in would be fabrication
carrying a designer's authority. A test asserts none of those strings
appear.
The IN-SEASON information architecture is NOT invented here. The spec
covers an offseason hub; nothing specifies how content, articles, wire and
the live slate share year-round navigation. That remains the open design
gap, and the route notes it.
Two test bugs caught and fixed: my first assertions matched my own doc
comments -- the ordering check found "WHAT CHANGED TODAY" in the header
block and the no-curves check caught the word "curve" in the sentence
explaining why curves are wrong. A guard that reads its own explanation is
not reading the render; both now strip comments first.
PHASE 3 — both surfaces registered under the reachability guard. Read-only
throughout, serving fingerprint unchanged, accrual clock unchanged at 0
eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — the order proposed E1 + E9 as Wave-1 and told me to follow the
audit if it disagreed. It disagrees: E9 calibration curve is WAVE D3,
gated on MODEL work ("resolve n>=20 vs N30, accrue buckets"), and E1
movement strip is WAVE D6, a large surface build. E9's gate is live right
now -- calibration is WITHDRAWN at 0 eligible dates, so the curve could
only render its empty state today. Building it would ship a component
whose entire purpose is unavailable.
PHASE 1 — three of the five real Wave D1 items were ALREADY DONE, and the
2026-07-31 audit has aged:
D1 glyph library audit: 38/83 wired (46%)
now: COMPLETE for everything wireable -- 39 of 83
designed glyphs map to a real archetype, all 39
are wired, colours match the registry exactly
(0 disagreements).
A1 card token audit: BUILT-BUT-DRIFTED, "in only 1 file"
now: BUILT-TO-SPEC -- it IS the --bg-1 token,
consumed by 32 files. The audit counted literal
hex, which is what a correctly tokenised value
looks like.
B1 boundary blue audit: PARTIAL, hex in 2 files
now: BUILT -- --priced-out/#8fb2de is a token with a
documented colour law, 4 consumers.
The 44 unwired glyphs are NOT a wiring gap: they have no backend
archetype, so wiring them means inventing 44 archetypes to consume
artwork -- the fabrication this programme refuses. That is the 41-vs-74
scope question and it is Kev's call. Separately, 2 registry archetypes
have NO designed glyph (DUAL THREAT, PAINT BOSS) -- a design gap.
PHASE 2 — what was genuinely absent is now built. web/src/lib/motion.js:
nudge() capped at 180ms so it reads as acknowledgement rather than
latency; bootStagger capped at 240ms because uncapped, row 40 waits 1.1s
and the stagger BECOMES the latency it exists to disguise; rowHover
returns handlers not CSS so touch cannot stick a hover state; and
revealOnIntersect returns an unobserve in every path and reveals
IMMEDIATELY when there is no IntersectionObserver or motion is reduced --
content is never hidden behind a capability check.
Reduced motion is honoured, not softened. The sharpest of the 10 tests:
bootStagger under reduced motion returns opacity 1, not merely delay 0 --
if the CSS animation supplies the opacity, skipping it leaves the row
invisible forever.
PHASE 3 — the reachability guard gains a PRIMITIVES section: a module
built to be embedded must declare its exports AND name its intended
consumers, because a primitive imported by nothing is the same
built-but-unread class as an unmounted component.
WAVE-2 UNGATED: F9-F11 offseason hub, F5 article media, E10/E12 Report,
E1 movement strip. GATED: E9 + E15 on model, E16/F8 on the resolution tail
and the card-system reconciliation, E2/E6 on licensing, E13 on another
order.
Read-only throughout; serving fingerprint unchanged; accrual clock
unchanged at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
READ-ONLY. Nothing designed, built or mounted.
THE HEADLINE: this inventory was largely already done. specs/
design-vs-build-gap-audit.md is a 61-item design-vs-build audit from
2026-07-31, every claim grep-verified, with a wave-ordered build plan. The
useful work is reconciling it, not redoing it.
WHAT ALREADY HAS A FULL DESIGN SPEC (and is NOT built):
- Offseason hub -- Vyndr Offseason.dc.html, 119KB: hub home desktop+390, an
NBA Summer-League variant with an OUTLOOK ONLY / NOT GRADED honesty
block, season-long board, season-read reveal with WHAT WOULD CHANGE THIS
READ, news/outlook feed row anatomy, quiet-wire empty state.
- Article media (S3) -- hero template plus four hero graphic archetypes,
all GENERATED data visuals never stock, inline figures with a one-pull-
quote max and a caption law that every figure names its data, article
card, OG 1200x630 with the master already rasterised.
- Movement strip (E1) -- steps not curves, green only when the move favours
the read, FLAT as hairline. THIS CORRECTS MY 2026-08-07 BOARD, which
listed it UNKNOWN and possibly satisfied by GradeShift. It is a specified
primitive that does not exist.
THE WIRE is BUILT (vyndr/Ticker on real ticker exhaust) and is a DIFFERENT
thing from NewsWire.tsx, which is the offseason news/outlook feed.
THE ONE REAL DESIGN GAP: the hub's information architecture across seasons.
The Offseason file specifies an OFFSEASON hub; nothing specifies what that
surface is IN-season, or how content, articles, wire and the live slate
share one navigation. Every component exists on paper; their composition
into a year-round media surface does not.
A FINDING AGAINST MY OWN RECENT WORK: I built a card renderer at 1080x1350
without checking whether a designed card system existed. It does -- five
master sizes with defined layouts and rasterised PNGs in exports/. The
content engine's cards are a parallel invention. Not wrong, but they should
conform to E16/F8 rather than diverge, and that is a design decision.
CORRECTION: the order lists content API, preview page, book-comparison
mount and guard widening as still pending. All four shipped at c575a70,
and book comparison was never unmounted -- my earlier board grepped only
web/src/app and missed component-level mounting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
two inventory errors
INVENTORY CORRECTION, and it was mine. Phase 2's two "orphans" are NOT
orphans -- my board grepped only web/src/app and missed component-level
mounting. The transitive check says both are already mounted:
BookComparisonPanel -> GradeResultCard -> app/scan/page.tsx
NewsWire -> ExploreHub -> app/explore/page.tsx
So book comparison is DONE (wired to /api/books, rendering on the grade
card) and THE WIRE is DONE-BY-DESIGN, mounted in ExploreHub. Its header
names an "Offseason Hub" as its home, and that hub genuinely does not
exist -- but that is board item #8, not a mounting bug, and inventing a
surface to satisfy a comment would be the wrong fix.
The lesson is the same one this session keeps teaching: I checked one
directory and reported a conclusion the check could not support.
ALSO CAUGHT: I overwrote src/routes/content.js, which was the Session-29
content-templates route, by picking a filename without looking. Restored
from git with no work lost; the new surface lives at
/api/content-studio and both now coexist.
PHASE 0/1 — /api/content-studio serves finished posts (copy, branded card,
card_svg, the fact_contract each was REQUIRED to have, and the facts that
actually backed it) plus a POST for editorial status in Redis. Private via
internal key; the Next proxy holds the key server-side so the browser
never does. /studio renders it as a thin client -- copy and card side by
side with the fact contract visible, because reviewing copy by reading it
is exactly how a wrong number ships. Never-blank: a night with nothing
generated says so.
API-FIRST is the point: the endpoint an autonomous poster will call is the
one the page already renders, so the agent handoff is a pointer change,
not a rebuild. Contract documented at docs/CONTENT-STUDIO-API.md.
EXPRESS 5 BROKE 23 SUITES at first: `router.get('/:date?')` throws at
mount time in Express 5, taking down everything that imports app.js. Two
explicit routes instead.
PHASE 3 — the reachability guard is widened from grade-fields-only to a
general built-but-unread check. Book comparison, THE WIRE and the content
studio are now registered surfaces; a page counts as its own entry point
(Next mounts it by convention) while everything else must trace to one.
22 checks green; a registered-but-unimported surface still goes red.
FULLY ISOLATED: read-only on model/slate/ledger, serving fingerprint
verified unchanged, accrual clock unchanged at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.
18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.
PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.
THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.
Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.
PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.
PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).
FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Inventory only, no build. Every line cites a file or a query; unverifiable
items are marked UNKNOWN rather than asserted.
Corrections to the assumed board:
- Scanner is DONE on the blue-boundary format (scan/page.tsx:715-717), not
still amber.
- Screen conversion is DONE except ONE file -- RouteStub survives only in
app/notifications/page.tsx across 47 route dirs.
- Terminal is deliberately retired (S57 redirect), not debt.
- ev_pct DOES render, in 5 frontend files. That flag was stale.
- edge_pct is NOT a scale bug: (model-line)/line is correct arithmetic that
explodes on small lines (a 5.5 projection on a 0.5 line is a legitimate
1000%). Live top values 900/860/700 on 68,364 rows. It is a display
defect, not a broken computation, and it reaches a user only via
alt_lines typing -- the grade card computes edge independently.
- All FOUR sports have real archetype registries (nba 15, mlb 15, soccer 6,
wnba 5) and ACTIVE_SPORTS runs all four. Only MLB BATTERS is MODELED;
everything else is scaffolding with a sound base rate and no proven
factors. MLB pitchers: pitcherEngine.js exists at 261 lines and is read
by ZERO serving code.
- Book comparison is built and mounted NOWHERE -- the same built-but-unread
class the reachability guard exists for, and NOT covered by it, since
that contract is grade fields only.
UNKNOWN and marked as such: opp_rank_stat prod population, internal-key
rotation, the ~9% void/DNP rate, and whether MovementStrip is a distinct
spec item or satisfied by GradeShift/MarketBreadth.
The ordered board puts five PARTIAL items first and flags accrual
sensitivity: seven items are isolated design/guard work buildable today
with zero clock impact; five touch the serving path and would muddy the
accruing repaired-champion dates. edge_pct is the sharpest tension -- a
real honesty defect a user can see, whose fix touches serving mid-accrual.
Clock today: 0 eligible dates, all four re-audit items WAITING.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — audit. espnStatsAdapter's slice(0,20) was already fixed at
494c83c, so the ESPN branch of getStatRows inherits a full log and is
CORRECT. MLB's l20_avg is a real seasonTotal/games aggregate, verified,
also CORRECT. TWO basketball defects remained:
DEFECTIVE nbaGameLogFeatures m20 = avg(vals.slice(0, 20))
-- l20_avg is the season reference projectionFor reads, and
slicing to 20 made it a twenty-game average wearing a season
label. The MLB l20 bug, unfixed for basketball.
DEFECTIVE getStatRows python getGameLogs(playerName, sp, 20)
PHASE 1 — both fixed via a named SEASON_LOG_DEPTH = 100, past an 82-game
season so a request can never truncate one. API COST: ZERO. The count is a
request parameter, so asking for a season is the same single call. No
extra request, no extra quota.
MEASUREMENT DEFERRED, EXPLICITLY: basketball is offline, there are no
settled basketball rows, and resolution before/after CANNOT be measured
now. This is a code fix, not a measured improvement -- exactly like the
deferred MLB sibling paths.
PHASE 2 — the window guard asserts the property on source: no fixed N may
stand in for a season. It immediately caught TWO MORE instances I had
missed in Phase 1 -- a hardcoded 20 at featureCache:377 and
gameLogService's own `count = 20` DEFAULT, which would have handed a
twenty-game window to any caller that omitted the argument. That is a
seventh path, found by the guard rather than by me.
Retro-confirmed: run unchanged against 981a05c it goes 3 failed / 6
passed, flagging the basketball slice, the game-log request and the
season-reference check.
The class in full, now six paths across two sports, every one of which
looked like ordinary code. `slice(0, 20)` is unremarkable; what made it a
defect was the QUESTION it answered -- "what is this player's season
rate?" -- and no test could see that mismatch because the value produced
was always a plausible number.
PHASE 3 — THIS CLOSES THE NON-ACCRUAL ARC. Everything buildable without
settled rows is built: push unblocked (it was the wrong remote, not the
firewall), grade surface honest and rendering, reachability guarded,
champion repaired across every path including dormant basketball, and the
window class guarded so it cannot return.
The program is now correctly IDLE on modeling. Today's count: 0 eligible
dates, all four re-audit items WAITING. FIRST TRIGGER: 10 eligible
calibration dates on the repaired champion, at which point the resumption
order is calibration re-fit, hits factor lift, prior verdict re-audit,
rbi lineup-slot gate.
No basketball measurement claimed. No NBA chain/archetype build. MLB
serving verified byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Three consecutive orders shipped a backend-correct field that never
reached a screen, all on a green suite: gradeBands (required by no
serving code), served_grade (dropped at the adapter boundary),
GradeScaleLegend (imported by nothing). Each was caught by luck on a later
re-check, and in two of the three I had already reported the wiring done.
WHY GREEN TESTS COULD NOT SEE IT: backend tests stop at the API payload.
They prove a field is PRODUCED and say nothing about whether it is
CONSUMED. Invisible by construction, not an oversight in any one test.
THE TRAP, NAMED: the difficulty pools in the backend, so by the time a
field exists on the payload it feels finished. What remains is a
three-line adapter change nobody considers worth verifying, so it gets
claimed rather than traced. The last inch is the one with no friction,
which is exactly why it gets skipped. "I added the field" and "a user can
see it" are different claims and only the first is fun.
THE GUARD traces each promised field the whole way: payload -> adapter
consumes -> component renders -> component is MOUNTED. Mounted is
transitive to a Next entry point (page/layout/template), the only thing
that puts a pixel on screen, depth-limited so an import cycle cannot hang
the suite.
Container rows are exempted EXPLICITLY, not silently: served_grade carries
container:true plus a rendersVia list, and a separate assertion checks
every named part actually renders. The exemption is auditable and cannot
hide an unrendered field.
The guard tests itself -- an orphan component must report unmounted, and
the contract must be non-empty, since an empty contract passing vacuously
is how this would most plausibly rot.
RETRO-PROOF: run unchanged against 3591c76, before the wiring, it goes
11 failed / 8 passed and names the exact bugs -- "the ceiling stance /
grade scale legend - its component is MOUNTED, not merely written", "the
served grade object - the adapter consumes it", "whether the band
separates from the baseline - a component actually renders it". Green on
the current tree.
HONEST SCOPE LIMIT: gradeBands is NOT in the contract and would not be
caught. It is a backend module, not a promised user-facing field, and it
is correctly unwired -- every band collapses to base-rate at current
resolution. Out of scope by design, not oversight.
Now in the standing suite, so the three-gate floor is tests green
(including reachability) + build exit 0 + fingerprint. No serving or model
change. p_win never mutated. No Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 re-check caught the same failure a THIRD time, mine again. Last
turn I created GradeScaleLegend.tsx and never mounted it, and I reported
that the separation flag "renders per band" -- it did not. grep: zero
frontend references to separates_from_base_rate, served_grade or
factor_adjustment. The honest grade was reaching the API payload and dying
at the adapter boundary.
That is three occurrences in three orders of the same shape: built,
correct, unread. gradeBands, then served_grade, then the legend.
PHASE 2 — the composed surface is wired end to end:
analyzeViaEngine1 -> served_grade + factor_adjustment on the payload
scan/page.tsx -> ScanResponse types them and forwards them
gradeAdapter -> gradeMeaning, separatesFromBaseRate,
bandRealizedRate, factorsApplied, refusalReason
GradeResultCard -> renders "WHAT THIS GRADE MEANS"
What a user now sees that they could not before: what the band has
actually realized, an amber note when the read CANNOT be separated from
the baseline, and -- only where a factor actually fired, with its proven
sign -- what moved the read, in plain language rather than feature names
("where he hits it vs who is standing there").
No narrative on props where nothing fired: factorsApplied is empty and the
block self-hides. Only the three PROVEN hits factors have labels, so an
unproven factor cannot acquire prose by being added to the map.
PHASE 1 — GradeScaleLegend is now MOUNTED on the grade card (compact). The
ceiling is a stated position where the grade is, not a page a user would
have to find.
PHASE 3 hand-verified through the real adapter across twelve states: B+
with 3/0 factors, B with 1, C+/C/C- all flagged not-separable, D, F, a
switch-hitter case where 2 of 3 factors fire, two refusals and a
no-forecast. never-blank PASS, no-manufactured-A PASS, no-narrative-when-
nothing-fired PASS, separation-flag-reaches-card PASS.
No A-threshold loosening. No calibrated number leaks. p_win never mutated.
engine_grade still read by zero serving code. No Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 caught my own repeat of the failure I diagnosed one order ago.
91927a4 attached `served_grade` BESIDE the old letter and left `grade`
alone -- so the honest grade reached nobody, exactly as gradeBands had
been built-correct-and-unread. grep showed served_grade appearing in one
file (where I set it) and all 14+ consumers -- scan route, dashboard,
parlay, newsletter, desk, content templates, retention -- still reading
`.grade`, i.e. still the dishonest letter.
CUTOVER IS NOW TOTAL: legacy.grade IS the honest letter. Overwriting the
one field every consumer already reads cuts every surface over at once
instead of editing fourteen call sites and missing one. engine1's index is
preserved as `engine_grade` and verified read by ZERO serving code.
Confidence follows the letter: it came from a grade-band midpoint of the
OLD letter, so leaving it would have paired a served B+ with a C's
confidence. Both now derive from p_win, kept on the existing 0-100 scale.
MEASURED BLAST RADIUS before shipping: 303 of 47,991 non-refused
snapshots (0.6%) have a grade but no p_win, and now render NO READ instead
of a letter. That is correct -- their old letter came from the retired
index carrying 0.48% resolution, i.e. noise -- and NO READ is a rendered
state with a reason, so never-blank holds.
PHASE 1 — the ceiling is now a STATED POSITION, not a confusing absence.
servedGrade.SCALE_LEGEND plus web GradeScaleLegend.tsx say it plainly: we
do not issue A grades, no band has hit at a rate that would justify one,
our honest ceiling is a strong B+ (~66% realized vs ~60% baseline), and if
the model earns an A the legend changes and we say why. The
separates_from_base_rate flag renders per band -- C+/C/C- are labelled
"we cannot separate this from the baseline", which is most of any slate.
PHASE 3 hand-verified across every state: B+ with 3 factors (basis
forecast_plus_matchup_factors), B+ with none (forecast_only), C flagged
not-separable, F, and three refusal states rendering NO READ with reasons.
never-blank PASS, no-manufactured-A PASS.
Test fallout was real and is documented rather than papered over: engine
BEHAVIOUR assertions moved to engine_grade, suppression assertions stayed
on grade (a suppressed prop has no letter either way), and the confidence
78 -> 95 change is the grade-band midpoint being replaced by p_win.
No A-threshold loosening. No calibrated number leaks (deployed set empty).
p_win never mutated. Ten frozen modules verified unchanged including
engine1 and probabilityEstimator. No Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
number beside it
PHASE 0 corrects the order's premise. A grade letter has been served all
along -- engine1.gradeProp builds it from an additive factor index,
computed INDEPENDENTLY of p_win. gradeBands is orphaned for a different
reason than assumed: it defines what a letter MEANS from realized
outcomes, and every band collapses to base-rate at current resolution.
The measurement that changed this order, on 3,417 settled props:
grade n realized mean p_win
A 8 0.500 0.647 <- the TOP grade did WORST
B 985 0.640 0.700
C 1,695 0.602 0.676
D 303 0.558 0.604
F 426 0.535 0.588
letter resolution 0.00116 (0.48% of variance)
p_win resolution 0.00715 (2.98%)
-> the letter carried 0.16x the information of the number beside it
Concretely, from the hand-verify: Christian Encarnacion's 0.95 over
graded C and his 0.05 under ALSO graded C -- same hitter, opposite
forecasts, same letter. The gap was never that grades don't ship; it is
that the weaker of two available signals shipped as the headline.
PHASE 1 — model/servedGrade.js derives the letter from p_win with bands
anchored on MEASURED realized rates (B+ 0.663 / B 0.646 / C+ 0.615 /
C 0.589 / C- 0.548 / D 0.512 / F 0.447, base 0.6005).
NO MANUFACTURED A, structurally: A+/A/A- are UNISSUABLE, not rare. The
realized rate plateaus at 0.65-0.68 above p_win 0.70, so no band has
earned a top letter; a test sweeps every p_win 0..1 and asserts none
produces one. Even 0.99 tops out at B+ with its realized 0.663 attached.
Raising that ceiling later is a deliberate, visible act.
Bands that cannot separate SAY so -- C+/C/C- carry
separates_from_base_rate false and copy naming it, which is the honest
description of a forecast explaining 3% of variance. Every grade states
its basis (forecast_only vs forecast_plus_matchup_factors, naming which
factors fired) and calibrated:false. engine1.grade is preserved as
engine_grade so nothing downstream breaks.
PHASE 2 — refusals render real states: insufficient_data -> "not enough
history to call this one"; juiced_no_edge -> "the book has priced the vig
past any edge on this side". 1,870 refused snapshots carry exactly those
two reasons and both now surface.
PHASE 3 — hand-verified on 12 real served props. Freeman/Rice/Encarnacion
0.95 overs now B+ (was B, C, B); the 0.05 unders now F (was C). Refused
doubles render NO READ with their reason. never-blank PASS,
no-manufactured-A PASS.
Serving change; nine frozen model modules unchanged including engine1;
p_win never mutated; no calibrated number leaks (deployed set empty); no
Bonferroni slot.
STILL TRUE: the forecast explains ~3% of outcome variance. This order did
not make the model better. It made the letter stop overstating it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The GATE-0 theory is REJECTED for VYNDR, tested rather than assumed. Both
remotes are ALREADY HTTPS -- there is no git@ SSH URL in this repo to be
blocked -- and both hosts answer on 443: git.builtbykev.com HTTP 200 in
0.64s, github.com HTTP 200 in 0.17s, Gitea's git endpoint 200, GitHub's
401 (auth required, reachable). No connectivity failure of any kind.
COLYRA's HTTPS-remote fix was right for COLYRA; VYNDR was already in the
state that fix produces, so applying it would have minted a new token to
solve a problem that did not exist -- and the pre-existing credential
would have made the new token look like the cure.
The real cause is in the error text, which named it exactly and was
misread all session: "could not read Username for 'https://github.com'".
Every attempt used `git push origin main`, and origin is GitHub
(kev3109/betonblk) with no stored credential. The WORKING remote is
`gitea` (git.builtbykev.com/builtbykev/vyndr.git), and a valid credential
for it sat in ~/.git-credentials the entire time. The habit of typing
`origin` is what kept twenty commits local. Nothing was blocked.
FIX: git push gitea main. Pushed 6452926..ecf78b9, 21 commits, verified by
git ls-remote matching local HEAD exactly.
No firewall rule touched, no token created, no code or model change --
infra only. The bundle and patch series stay as belt-and-braces; they are
no longer the only copy.
docs/GIT-PUSH-DIAGNOSIS.md records it so no future session re-reaches for
an infrastructure theory when the error message names a credential and a
specific host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 1 — the existence risk is resolved. Nineteen commits were local-only
with no push credentials. Three artifacts now exist off the working tree:
~/vyndr-full-history-2026-08-07.bundle 6.8M SELF-CONTAINED, clone it
~/vyndr-session-2026-08-07.bundle 244K needs existing history
~/vyndr-session-patches/ (20 patches) 1.5M
Use the full-history bundle -- `git bundle verify` shows the session bundle
requires ref 6452926, so it only applies onto a repo that already has this
history. The full one clones standalone. Copying one off the machine is
the top non-accrual action item.
PHASE 2 — the watch, today's starting line: 0 eligible rows, 0 eligible
dates, everything WAITING. Zero is correct, since the repair ships in this
session's commits and no settled row can carry the marker yet.
THRESHOLDS ARE ATTEMPT FLOORS, NOT TRUST FLOORS, and this is written into
the ledger so a future session cannot misread it. Reaching 10 dates means
the calibration re-fit CAN BE MEASURED, not that it is trustworthy -- we
lived that distinction tonight with 2-4 cluster CIs, a LODO gate at
1.4-9.3% power, and a >=40 bar that was right for promotion and wrong for
deploy. A 10-date map is thin, deploys PROVISIONAL with auto-demotion, and
its interval will still be wide.
Real-time estimates stated so the wait reads as designed: ~2 weeks for
calibration and hits lift, 3+ weeks for the verdict re-audit and the rbi
gate, since rbi accrues slowest and a date only counts once its props
settle.
PHASE 3 — resumption order fixed, triggered purely by date thresholds:
calibration re-fit @10, hits composed-lift re-measure @10 (the 1.39%
figure is VOID, direction unknown), prior verdict re-audit @14 (not
pre-priced), rbi lineup-slot two-part gate @14. FIRST TRIGGER: when
reAuditEligibility reports 10 eligible calibration dates.
Until then the program is honestly IDLE on modeling. That is the correct
state, not a gap to fill -- the alternative is measuring on reconstructions
of a retired forecast, which this program has now refused by name three
times.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
repair falsified
PHASE 0 — four checks PASS, one defect found and fixed.
PASS gradeBands is required by NO serving code -- built across several
orders, never wired. No stale band derived from the retired
ten-game champion can reach a user, because none reaches a user
at all.
PASS CALIBRATION_DEPLOYED is [] and the calibrate loop iterates it, so
calibrate() is never called and p_win_calibrated is never set. The
only assignment site sits inside that empty loop. No withdrawn map
can leak.
PASS projectionFor reads l20_avg, which mlbGameLogFeatures now builds
from fullLog -- so refusals are computed on the repaired
full-window reference, not the retired ten-game one.
PASS factors still fire with correct sign across the repaired base
range (0.35/0.50/0.65/0.80): defense lowers, pitcher-contact
raises, platoon raises at every point. Mechanical firing check
only -- NOT a lift re-measurement, which waits for accrual.
DEFECT FIXED — my own repair falsified a user-facing sentence. The grade
card rendered "Last 20 games average: X" from l20_avg, and l20_avg is now
a FULL SEASON average. The number changed and the label did not, so the
surface was stating something the data no longer supported. Copy now reads
"Season average"; trapDetection's L20 explanations likewise. The field
name is kept -- it is read in many places -- but no rendered sentence
claims a window that isn't there.
That is the same class as everything else tonight, one layer out: a
correct-looking string describing data that moved underneath it.
Serving change; frozen model modules unchanged; p_win never mutated; no
Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
in code
PHASE 0 — getStatRows is the single base-rate path, so every branch is
audited, plus the feature builders since l20_avg is the season reference
projectionFor reads:
getStatRows MLB -> estimator base fullLog CORRECT (929fd81)
mlbGameLogFeatures l5/l10/l20 last10 = 10 DEFECTIVE
espnStatsAdapter.parseGameLog slice(0,20) DEFECTIVE
getStatRows NBA/WNBA ESPN branch inherits 20-cap DEFECTIVE via source
getStatRows NBA/WNBA python branch getGameLogs(...,20) dormant (offline)
pitcherEngine / skillProjection statcast profiles N/A
pitcher props via getStatRows MLB fullLog CORRECT
settleSource full log (S64) CORRECT
THE PITCHER ANSWER IS GOOD NEWS: pitcher props run through the same
getStatRows MLB branch, so 929fd81 repaired them too. There is no separate
defective pitcher base-rate path.
THE ONE HIDING IN PLAIN SIGHT: mlbGameLogFeatures carries the comment
"l20 = all available (the season per-game reference projectionFor needs)"
while building from last10 -- so l20_avg was a TEN-GAME AVERAGE WEARING A
SEASON LABEL, feeding both the consistency pull inside the estimator and
projectionFor, which decides refusals. It survived the previous repair
because that fix touched only getStatRows.
PHASE 1 — mlbGameLogFeatures now reads fullLog; espnStatsAdapter drops its
slice(0,20) cap. ZERO new API calls on both: each widens data already
fetched and then discarded, the same shape as the original repair. The
python branch is left alone -- the service is offline in prod and fixing it
would be speculative.
Their before/after resolution is NOT measured, deliberately: the only way
to measure today is to reconstruct the repaired forecast over old rows,
which is the reconstruction-vs-served trap this order refuses. Code fix
now, measurement at accrual.
PHASE 2 — MODEL_VERSION bumped to engine1@2026-08-07-fullwindow, so every
forward snapshot is self-identifying (retentionService already stamps it;
no new plumbing). model/reAuditEligibility.js encodes the rule: isEligible
accepts only the repaired marker, assess counts eligible DATES not rows,
and ACCRUAL is frozen at calibration 10 / hits-lift 10 / verdict-reaudit
14 / rbi-gate 14. A test locks the invisible case -- a MIXED table of 330
rows with 30 repaired returns eligible_dates 3, not 330 rows of false
confidence. Once both generations share a table a naive count would fit a
map on a blend of two forecasters.
PHASE 3 — the board, each consequence labelled: calibration WITHDRAWN
(refits at 10 dates, never on reconstructions); factor verdicts SUSPECT
(all measured against a champion worse than a frequency table, direction
UNKNOWN, not pre-priced, 14 dates); hits factor lift UN-REMEASURABLE (10
dates, factors still wired and transmitting); rbi lineup-slot RE-QUEUED
(14 dates). Pre-registered order: calibration, hits lift, verdict
re-audit, rbi gate.
Then STOP and accrue. Nothing further can be honestly measured until the
board fills with rows the repaired champion produced.
Serving-path changes by design for the MLB feature path and NBA/WNBA logs;
eleven frozen model modules verified unchanged. p_win never mutated. No
Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time
baseline (each player's rate over games strictly before that date, >=10
prior games, box scores back to 05-01), the served champion LOSES on all
four stats, three of four CIs excluding zero:
hits 0.00251 vs 0.00774 CI [-0.0074,-0.0011]
TB 0.00393 vs 0.00619 CI [-0.0055,-0.0003]
rbi 0.02481 vs 0.03133 CI [-0.0153,-0.0005]
runs 0.00181 vs 0.00683 CI [-0.0114,+0.0008]
PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability
builds its base rate as the frequency over every row it is handed, and
featureCache.getStatRows handed it res.last10. So the "season rate" was a
TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The
0.40 recency weight costs resolution on all four stats (-0.00086,
-0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on
hits and rbi, marginally helpful on TB and runs -- so they are left alone.
PHASE 2 — two lines, no new data, no extra API call, because fullLog was
already fetched by the same adapter call that produced last10:
getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20.
hits 0.00251 -> 0.00817 (tripled; now above the fair baseline)
TB 0.00393 -> 0.00734 (above baseline; vs old CI [0.0020,0.0067])
rbi 0.02481 -> 0.02727 (still below baseline, CI includes zero)
runs 0.00181 -> 0.00436 (still below baseline, CI includes zero)
Gate stated exactly: hits and TB now exceed the fair baseline on the point
estimate; rbi and runs remain below but EVERY CI now includes zero, so no
stat reliably loses to a frequency table. That is a tie on rbi/runs, not a
win, and it is reported as one. Only TB's improvement over the old
champion is CI-confirmed; the rest are directional.
STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were
fitted on the retired forecast and fromLedger cannot rescue them -- settled
ledger rows still carry OLD p_win, so refitting today would refit the
retired forecast. Nothing is served calibrated until dates settle under
the repaired champion, and the favourite-longshot bias must be re-measured
rather than assumed to survive. The shadow duel is void.
PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it
needs settled rows produced BY the repaired champion, which ships in this
commit. Replaying would score the factors against a reconstruction rather
than the served forecast. Deferred, explicitly. The factors remain wired
and transmitting; only their lift is unquantified on the new baseline.
PHASE 4 — standing flag, and it is large: EVERY factor verdict in this
programme, every null and every THEATER, was measured against a champion
worse than a frequency table. Signal added to noise reads as noise. Prior
verdicts may deserve re-audit. Logged, not re-run.
Re-queued not built: rbi lineup-slot / RISP opportunity through the
two-part gate, now landing on a repaired champion.
Serving-path change by design; the byte-identical invariant inverted and
all four stats move. Nine frozen model modules verified unchanged. No
Bonferroni slot -- resolution accounting on the champion's own knobs.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
out-resolved by a frequency table on three of four stats
PHASE 0 — the 14.51% is REAL. Re-derived with a paged pull asserted
against an exact count (rbi 7,930 == 7,930; hits 11,690; TB 12,086; runs
6,440), since this harness produced a false null three times tonight. rbi
resolution 0.03268 reproduces, deciles are monotone through the middle,
and 20 raw rows are in the artifact for hand audit.
CAVEAT GOVERNING EVERYTHING BELOW: the naive forecasts are leave-one-out
ON THE EVALUATION WINDOW, so they see the rows they are scored on while
the model is strictly point-in-time. They are upper bounds on available
resolution, not fair competitors, and every comparison is read that way.
PHASE 1 — the split:
stat MODEL (a)player-base (b)lineup-slot (c)within-stratum
rbi 0.03268 0.01167 0.03608 0.01908
hits 0.00252 0.00446 0.00100 0.00473
TB 0.00442 0.01331 0.03448 0.00607
runs 0.00130 0.00262 0.01156 0.01170
FINDING 1 — rbi's resolution is LINEUP ROLE almost exactly. Batting-order
slot alone resolves 0.03608 against the model's 0.03268. A single integer
accounts for the whole anomaly and slightly more. That is opportunity, not
skill -- the cleanup hitter bats with runners on. 36% is matched by player
identity alone. Within similar-base-rate strata the model still resolves
0.01908, 58% of its total and higher than any other stat's ENTIRE model
resolution, so genuine within-role discrimination exists on top.
FINDING 2 — on three of four stats the model is beaten by "he's a .270
hitter". Base-rate-only out-resolves the model 1.8x on hits, 3.0x on TB,
2.0x on runs. Even allowing for the window-peeking advantage, a 1.8-3.0x
gap is not explained by that alone: the served counter appears to DESTROY
discrimination relative to the player's own rate. rbi is the one stat
where the model beats the naive baseline.
FINDING 3 — lineup slot out-resolves the MODEL on three stats: TB 7.8x,
runs 8.9x, rbi 1.1x. Hits is the only stat where batting order carries
less, which is mechanically right -- a hit is a hit wherever you bat, but
runs, RBI and total bases all scale with opportunity.
PHASE 2 — all three worlds are partly true, in measured proportions.
World A ~90% true (slot covers rbi's entire resolution). World B ~36% true
for rbi, but the WHOLE story for hits/TB/runs where base rate alone wins.
World C true with a low ceiling: hits' total available spread resolution
is 0.00446, i.e. 1.8% of variance from a forecast that has seen the
answers.
PHASE 3 — the next arc is NOT "strengthen hits factors". Hits has the
lowest available resolution on the board and last order's wiring already
took it to 1.39% of a ~1.8% ceiling. Named first factor order for next
session: LINEUP SLOT / RISP OPPORTUNITY on rbi through the two-part gate --
input already ingested and prod-verified (S89), resolution measured not
hypothesised, causally-correct unit is plate appearances with runners on.
Measured availability is not a pass; it still faces the gate.
And higher-value than either: the counter being out-resolved by a
frequency table on three of four stats is a defect in the CHAMPION, not a
factor problem, and it costs nothing to test -- the recency blend and the
+/-0.03 / +/-0.015 nudges are three lines in probabilityEstimator.
The hits transmission win from 43f65d3 stands: the conduit is real and
permanent. This order changes only which stat has the most worth flowing
through it.
Diagnostic only -- no factor wired, no serving path changed, p_win
untouched, all frozen modules byte-identical. No Bonferroni slot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
inconclusive
THE BUG THIS NEARLY SHIPPED AS A FINDING. The first audit reported 0
factors fired on all 1,140 rows. Not a result -- my paging helper ordered
by `id`, and batter_spray, team_defense, platoon_splits and
statcast_aggregates have composite primary keys with NO id column. The
query errored, the loop broke on error, and four fully-populated tables
read as empty. hitsFactorContext.js -- the PRODUCTION loader -- had the
identical defect, so live wiring would have loaded nothing and served
unadjusted while logging success. Third occurrence of this class in one
session. Both loaders now order by a real column and THROW rather than
degrade. The Phase 2 gate is what caught it: no resolution number was
quoted until transmission was proved.
PHASE 1 — pipeline is now base -> FACTORS -> CALIBRATE -> GRADE. Context
built in snapshotService BEFORE gradeAndCacheSlate (was line 640+, grade
at 454), threaded per prop, applied to p_over before p_win is set with
p_win_prefactor and a full trace retained. Hits only. Coverage 859/1140
rows (75%): 474 with all three factors, 256 two, 129 one, 281 none.
PHASE 2 — TRANSMISSION PROVEN, 12/12 sign-correct, 4/4 per factor, each
applied IN ISOLATION. My first table compared each factor's expected sign
against the COMPOSITE change and showed 3 false failures -- with three
factors firing the net can oppose any single member; that was a flaw in
the test, not the wiring. Two under-side rows confirm the flip is handled:
a factor raising p(over) correctly lowers p_win. Switch hitters (Bailey,
Bell, Rocchio) took no spray adjustment while their other factors fired
normally -- the refusal is selective, not a blanket skip.
PHASE 3/4 — both maps refit on the factor-adjusted forecast; the
shadow-duel baseline is VOID and restarts, since it accumulated against a
different forecast. Point-in-time, 765 held-out rows:
reliability 0.00795 -> 0.00828
RESOLUTION 0.00229 -> 0.00345 (variance explained 0.93% -> 1.39%)
Brier 0.25398 -> 0.25305 delta -0.00093 CI [-0.00225,+0.00002]
Resolution rose 51% relative. The CI TOUCHES ZERO on 4 eval dates, so the
composition does NOT earn a proven keep -- three isolated passes did not
grant a composed pass. INCONCLUSIVE, reported as such. The gain is far
below the sum of the isolated effects, which is expected: all three run
through the same pitcher-batter confrontation and share signal.
PHASE 5 — 1.39% of variance is still far below what band separation
needs. The pivot was correct and incomplete: the plumbing defect was real
and is fixed, three proven factors reach the served number for the first
time, and transmission alone did not buy grade separation. Next arc is
factor STRENGTH and BREADTH, not more plumbing.
PHASE 6 — rbi anomaly logged, not chased: 14.51% variance explained vs
hits 1.03%, on the stat we do not serve corrected and which has no proven
factors. Either the biggest lever on the board or a mirage; it deserves
its own order.
The byte-identical invariant INVERTED for hits by design. All 13 frozen
non-hits modules verified unchanged, probabilityEstimator included -- the
factors ride outside it. No new Bonferroni slot; the composed OOS claim is
reported with its CI and not claimed as a pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
— the proven factors were never wired in
PHASE 0 — two truths recorded. The swap is a BET, not an OOS win:
isotonic beat low-param on identical held-out rows (hits +0.0028, rbi
+0.0042, TB tied) and we serve low-param anyway on an untestable prior
about shared daily structure. At 19 dates nothing here can test it. And
the MIN_SLOPE catch is preserved as standing rationale: a near-zero or
negative slope collapses toward base-rate-for-everything, which LOWERS
Brier while destroying all resolution -- a metric win that guts the
product.
PHASE 1 — the duel is now falsifiable. Both corrections computed on every
hits/TB prop; p_win_lowparam served, p_win_isotonic_shadow logged in its
own try so it can never break serving. calibrationDuel.adjudicate encodes
the rule IN CODE before any forward date exists: >=10 forward dates and
isotonic winning with a date-block CI excluding zero => REFUTED, revert;
otherwise UPHELD; under 10 dates PENDING regardless of the numbers. A
date counts as forward only if NEITHER map was fitted on it -- otherwise
we would be scoring which map memorised better. Nothing swaps now.
PHASE 2 — the ceiling, quantified via Murphy decomposition:
stat reliability RESOLUTION uncertainty variance explained
hits 0.01353 0.00252 0.24532 1.03%
TB 0.01419 0.00442 0.24329 1.82%
rbi 0.00654 0.03268 0.22531 14.51%
runs 0.00788 0.00130 0.23182 0.56%
Calibration did exactly what theory says and nothing more: hits
reliability 0.01353 -> 0.00233 (-0.0112, 83% of the error removed) while
resolution moved -0.0002. Unexpected: rbi has 13x the resolution of hits
and is the one stat we do NOT serve corrected -- it needs calibration
least and discriminates most.
PHASE 2 DIAGNOSIS — NOT-TRANSMITTED, and not weak, ABSENT. Traced in code:
sprayDefense.js and platoonSeverity.js are required by NOTHING in src/,
only by analysis scripts and their own tests. The served p_win
(intelligence/probabilityEstimator.js:54) reads exactly four inputs --
game-log frequency, opp_rank_stat +/-0.03, home_away +/-0.015, and a cv
pull -- with zero occurrences of spray, platoon, hard-hit or
contact-profile. And snapshotService grades at line 454 while computing
challenger/context at 640+, so everything proven is computed DOWNSTREAM of
the grade it would inform. The three proven hits factors have never once
moved a served number.
That reframes the recent nulls: "calibrated p_win does not separate within
archetype" was never a statement about factors. The factors were not in
the forecast.
PHASE 3 — bands rebuilt on SERVED values (hits/TB low-param, rbi/runs
raw): 28 archetype slots across four stats, ZERO show lift. No longer an
open shrug -- it is the arithmetic of resolution 0.0013-0.0327 against
uncertainty ~0.23. A forecast explaining 1% of variance cannot produce
separating bands, and no correction to its numbers will change that.
HEADLINE: calibration is complete, delivered honest numbers on two stats
and zero grade separation, because the counter has no resolution -- and
the proven factors are not wired into the forecast at all. The second is
the reason for the first, and it is plumbing rather than a modelling wall.
Per-archetype grades need proven factors that actually reach p_win. Last
calibration order.
Serving unchanged from 74cf1ce. p_win never mutated. No Bonferroni slot.
Counter and frozen clusters verified file-by-file (15 modules).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — sample-limit truth on record: on 19 dates BOTH stability
instruments are underpowered. LODO power 0.014-0.093 (best 0.337 across
every k tried); deploy CIs rest on 2-4 date clusters, where a
cluster-robust interval has ~1 df. This is the SAMPLE, not a fixable
instrument, and the gate-refinement loop stops here. Runs corrected: its
DATE-DRIVEN label was an artefact of the coin-flip ruler (2 reversals in
3 drops never cleared cutoff 2) -- it is an ordinary no-fittable-map
refusal.
PHASE 1 — the bias is ROBUST, tested model-free and map-free with a
date-block bootstrap. Pooled over-prediction rises monotonically -0.0076
/ +0.0428 / +0.0963 / +0.1589 / +0.2451 across deciles from 0.5 to 1.0,
sign stability 0.9946 over 17 date blocks, and 4 of 4 stats replicate
(bar was 3). Also visible: realized rate PLATEAUS at 0.65-0.68 from p=0.7
upward -- the 0.9+ bucket (0.6624) does no better than the 0.8-0.9 bucket
(0.6841). The model has no high-confidence reads, only high-confidence
numbers.
PHASE 3 — Platt, two parameters over the whole curve, shrunk toward
identity by fit-date count. Validated as a NEW estimator vs RAW with
date-block CIs:
hits a=0.406 shrink 0.565 0.2626 -> 0.2540 CI [-0.0112,-0.0069] DEPLOY
total_bases a=0.472 shrink 0.333 0.2490 -> 0.2429 CI [-0.0062,-0.0059] DEPLOY
rbi a=0.775 shrink 0.231 0.2011 -> 0.2007 CI [-0.0007, 0] REFUSE
runs a=-0.032 REFUSE
A GUARD THE FIRST RUN NEEDED: runs fitted a = -0.032. A non-positive
slope inverts the forecast rather than flattening it, and near zero the
curve collapses to a constant predicting the base rate for everything --
which LOWERS Brier while destroying all resolution. It would have scored
as a win while making the product worthless. MIN_SLOPE now refuses it by
name, with a test.
STATED PLAINLY: on the identical held-out rows isotonic BEAT the
low-param on hits (+0.0028) and rbi (+0.0042) and tied on TB. The swap is
a CAPACITY JUDGEMENT, not a measurement -- the window spans 2-4 date
blocks and that is exactly what a flexible map produces when it captures
structure shared by fit and eval. Labelled as a judgement.
PHASE 4 — hits and total_bases serve the correction, basis
direction_robust_magnitude_provisional (direction bootstrap-robust,
magnitude thin-sample and shrunk). rbi is WITHDRAWN to raw -- it was
deployed on isotonic at ced4042 and the low-param does not beat raw.
runs stays raw. Auto-demotion still armed.
PHASE 5 — the standing finding, stated hard: across 18 archetype slots on
three stats, calibrated p_win separates within archetype NO BETTER than
raw. Every slot is one band indistinguishable from its base rate, zero
show lift. Per-archetype separation is not coming from calibration; it
comes from proven factors or it does not exist. Five orders of
calibration have delivered what they can -- honest numbers on two stats --
and nothing on the question the grade product turns on.
p_win never mutated; no Bonferroni slot; the robust-claim test ran before
any calibrator was built and could have ended the session at Phase 2.
Counter and frozen clusters verified file-by-file, including calibration.js
and calibrationService.js, both untouched and simply off the serving path.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
FAILs were false
PHASE 0 — the gate at 1f40014 was mine and was an incoherent pair. A 1-SE
informativeness bar with a ZERO-reversal rule: at exactly 1 SE a stable
stat's drop reverses with prob Phi(-1)=0.1587, so on four informative
drops P(>=1 reversal | perfectly stable) = 1 - 0.8413^4 = 0.50. It failed
stable stats half the time by construction. And the pooled n*=70
mis-credited EVERY stat -- too low for hits (own 77) and runs (81), too
high for total_bases (60) and rbi (54).
PHASE 1, blind. Per-stat (g, sigma_row): hits -0.01288/0.11251, TB
-0.01380/0.10680, rbi -0.00884/0.06459, runs -0.00902/0.08080. All four
clear z=1.96 at full n, so none is NO-EFFECT. Committed k=1 with per-stat
n* and a binomial cutoff holding FP at 0.004-0.031.
THE FINDING THAT DOMINATES: the test has no power. Against a strong
instability (date-to-date SD equal to the effect) it detects a failure
1.4%-9.3% of the time, and across every k from 1.0 to 2.0 the best any
stat reaches is 0.337. A gate that cannot fail cannot pass, so
LODO_POWER_FLOOR=0.50 makes UNTESTABLE structural -- "could not test" can
never read as "passed".
PHASE 2/3 cold, at each stat's OWN n*:
hits 5 informative, 0 reversals, cutoff 2, power 0.093 UNTESTABLE
TB 5 informative, 0 reversals, cutoff 2, power 0.093 UNTESTABLE
rbi 4 informative, 1 reversal, cutoff 2, power 0.045 UNTESTABLE
runs 3 informative, 2 reversals, cutoff 2, power 0.014 UNTESTABLE
Setting the power floor aside entirely, NOT ONE STAT EXCEEDS ITS CUTOFF.
PHASE 4 — rbi's FAIL was false, as the order suspected. So was RUNS' --
which the order did not anticipate, having classified it DATE-DRIVEN on a
244-row reversal; two reversals in three drops does not clear a cutoff of
2. TB's PASS was vacuous: the test could not have failed it. hits' own n*
is LARGER than the pooled one (77 vs 70), and it remains untestable.
PHASE 5 — deploy basis is now the date-clustered CI alone:
hits CI [-0.0139,-0.0097], 4 date clusters relabelled ci_only
TB CI [-0.0061,-0.0045], 2 date clusters RELABELLED, kept
rbi CI [-0.0092,-0.0010], 2 date clusters NEWLY DEPLOYED
runs no fittable map at its split REFUSE, no CI either
Every deployed stat carries calibration_basis ci_only_lodo_untestable and
auto-demotion is the SOLE stability guard, not a backstop to a passed
test. Stated plainly: those intervals rest on 2-4 date clusters, which is
thin, and it is now the only support. rbi gains chainAcross stackability;
its bands rebuilt on p_win_calibrated (425 rows) are every-archetype
base_rate. runs is queued for the low-param calibrator for the ordinary
reason -- no fittable map -- not on the date-driven finding, which was an
artefact.
PHASE 6 — the deploy set was set by a coin-flip-power ruler; it is now set
by a per-stat power-coherent pre-committed test whose first act was to
report that it cannot evaluate anything. The audit was permitted to wound
the live deploy and did: total_bases lost its LODO claim. Standing
question unchanged -- 18 archetype slots across three deployed stats, every
one a single band indistinguishable from base rate.
Blind ordering held. p_win never mutated. No Bonferroni slot. Counter and
frozen clusters verified file-by-file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
routed as date-driven
PHASE 0 — threshold derived BLIND, before any stat was re-read. A
reversal is informative only if that date's Brier delta is
distinguishable from zero at its row count. Per-row Brier difference
d_i = (pc-y)^2 - (p-y)^2, so SE(n) = SD(d)/sqrt(n) and
n* = (SD(d)/|effect|)^2. Pooled across all four stats so no single
stat's verdict could shape the threshold deciding it:
pooled rows 3,417 | SD(d) 0.09816 | |effect| 0.01175
n* = (0.09816/0.01175)^2 = 69.8 -> 70
The hand-chosen 20 sat at 0.54 SE -- a coin flip. That is the defect
this removes, and why the previous verdict moved with the number.
Committed as calibrationRegistry.LODO_MIN_HELD_ROWS = 70 with
LODO_THRESHOLD_BASIS; a test recomputes (SD/effect)^2 and asserts it
equals the constant, so it cannot drift from its own justification. The
derivation script prints no stat verdict, no date and no reversal.
PHASE 1 — LODO at n*, applied cold:
hits 5 informative drops, 0 reversals PASS
total_bases 4 informative drops, 0 reversals PASS
rbi reverses 2026-08-01 (n=99) FAIL
runs reverses 08-01 (n=86), 08-05 (244) FAIL
hits held-out deltas -0.0041/-0.0080/-0.0192/-0.0140/-0.0139 across
123-272 row dates, favourite sign holding on every testable drop. THIS IS
THE INSTRUMENT FINALLY POWERED, NOT VINDICATION OF A PREDICTION -- the
withdrawal at 6ae11f1 was correct on the instrument available then, which
admitted 20- and 25-row dates as evidence. Nothing about hits changed;
the threshold stopped being chosen.
PHASE 2 — both failures are DATE-DRIVEN, not underpowered. Every
reversal sits above n*=70 (99, 86, 244), so no threshold and no further
accrual rescues either: isotonic is fitting day-structure. Routed to the
low-parameter calibrator queue (Platt/beta), not built here.
PHASE 3 — CALIBRATION_DEPLOYED is now ['hits','total_bases'], frozen and
tested, both PROVISIONAL with auto-demotion armed and the >=40
date-cluster promotion bar unchanged. hits stackability for
chain.chainAcross is RESTORED, and the record shows it returned through
the powered gate rather than by fiat. hits bands rebuilt on
p_win_calibrated (765 eval rows): every archetype still one band, still
base_rate -- calibrated YES, proven-per-archetype NO.
PHASE 4 logged: the deploy set is now set by a power-derived,
pre-committed, tested constant rather than an operator-chosen number. At
6ae11f1 that rule moved the live path AGAINST the operator; it has now
moved it back on the same evidence because the instrument changed. Both
directions are the rule working. And calibrated p_win separates within
archetype no better than raw across 13 archetype slots on two deployed
stats -- per-archetype separation will come from proven factors or not at
all.
p_win never mutated; no Bonferroni slot consumed; counter and frozen
clusters verified byte-identical file by file.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PHASE 0 — I applied factorGate's >=40 date-cluster floor to a calibration
layer without challenging the binding. That floor is a cluster-robust
interval bar for a CAUSAL claim. Calibration makes no causal claim, has a
bounded failure mode (it can only over- or under-shrink) and consumes no
Bonferroni slot. Its real risk is that the correction is DATE-DRIVEN, and
leave-one-date-out tests that directly -- a STRICTER bar, since a cluster
count cannot detect a single day carrying the effect. The >=40 floor is
retained, correctly scoped as the PROMOTION bar.
PHASE 1 — both guards codified, 11 tests, green before Phase 2.
Demonstrated on live data: raw population violated=true, mean_p 0.4962,
both_sides_share 0.9763; after dedup violated=false, mean_p 0.6694. The
null guard's test demonstrates the trap explicitly, since (null-1)**2 is
1 and (null-0)**2 is 0 so a Brier over nulls equals the win rate.
PHASE 2 — LODO:
hits n=1140 dates=17 2 reversals (07-22 n=20, 07-26 n=25) FAIL
total_bases n=1050 dates=7 0 reversals, 0 sign flips PASS
rbi n= 630 dates=5 1 reversal (08-01 n=99) FAIL
runs n= 597 dates=5 2 reversals (08-01 n=86, 08-05 n=244) FAIL
Threshold sensitivity reported because the verdict moves: total_bases
passes at every held-size threshold, runs fails at every one, and hits
fails ONLY when 20/25-row dates are admitted. I fixed MIN_HELD_ROWS=20
before seeing which stats passed and did not move it afterwards to
preserve a deploy. Honest caveat: a per-date Brier delta on 20 rows has a
standard error several times the effect, so the instrument is
underpowered per-drop -- an argument for pre-registering a higher
threshold, which is a Roundtable call, not one to make while holding the
results.
PHASE 3 — total_bases DEPLOY-PROVISIONAL, band [0.6-0.8]. hits, rbi and
runs REFUSE.
HITS WAS BEING SERVED CALIBRATED AND IS NOT ANY MORE. snapshotService
hardcoded it since S91; it fails LODO, so it is out. A stat that cannot
survive dropping one day was never calibrated, it was fitted to that day.
The consequence is real -- hits props become unstackable for
chain.chainAcross -- and it errs toward withdrawing a claim rather than
preserving one on a fragile verdict. Deployment is now driven by a frozen,
tested CALIBRATION_DEPLOYED set, not a hardcoded stat name.
PHASE 4 — calibrationRegistry, 14 tests. Deploy needs BOTH gates, neither
waivable. reverify auto-demotes on the first breach (CI stops excluding
zero, or the favourite bias flips sign) and logs the breaking date.
Promotion needs the original >=40 bar. A provisional deploy that cannot be
taken away is just a deploy.
PHASE 5 — TB bands rebuilt on calibrated values, 625 eval rows. The
two-bar rule still bites: calibrated YES, proven NO, so they stay a
base-rate read, now honestly numbered. Every archetype still collapses to
one band -- calibrated p_win separates within archetype no better than raw.
PHASE 6 logged only: the dead gradient is buried (hits~TB > runs > RBI,
and RBI has the SMALLEST bias, so the skill-driven-gradient mechanism did
not survive); the refused set is a map of missing inputs; a low-parameter
calibrator is queued unbuilt.
p_win never mutated; calibration rides as p_win_calibrated with
calibration_status provisional. No Bonferroni slot consumed. Counter and
frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Settlement done (15,484 written). Calibration improves held-out Brier on
all three stats it can be fitted for, beating every factor ever tested.
No stat deploys: the date-cluster ceiling is 17, not 90.
PHASE 0 CORRECTIONS: 71,192 snapshots unsettled, not 22,032. Span is
07-19 -> 08-06 = 19 dates, not 05-01 -> 08-04. Nothing has ever been
rescaled on any stat -- all four are base-rate bands today -- and the TB
"inversion confirmed" was the units-bug artifact, UNPROVEN.
PHASE 1, two integrity findings both caught by the gate:
1. The dupe check hard-failed on snapshot id 33875. model_snapshots is
written by the cron at 14/19/22/1/3 UTC and an unordered .range() walk
over a live table returns overlapping pages. Fixed with .order('id').
2. 12,894 rows were logged AFTER first pitch -- cycles at ET 21/22/23 on
the game date (10,738) plus 664 the next morning. A 01:00-UTC cycle is
21:00 the previous evening Eastern, same game date, two hours into the
slate. Tested for contamination: bias +0.0058 in-game vs +0.0008
pre-game, so NOT sharper, just late. Excluded for provenance.
THE ENABLING MOVE DID NOT ENABLE. 71,192 rows collapse to 4,799 distinct
pre-game props (2.5x cycle fan-out, then 97.6% both-sides duplication,
then the pre-game filter). Hits ends at 1,140 rows against the ledger's
existing 1,312. Date-clusters: hits 17, TB 7, rbi 5, runs 5.
THE MEASUREMENT THAT NEARLY WENT THE OTHER WAY: 97.6% of props carry both
sides, whose p_wins sum to ~1 and whose outcomes are complementary, so
the raw population is pinned to 0.5 by construction. Measured that way
the counter reads +0.0002 on hits -- "perfectly calibrated" -- and would
have overturned three sessions. Deduped to the model-picked side it is
+0.0868. The tell was mean p_win sitting at 0.4998 on every stat.
PHASE 2/3, isotonic point-in-time, split by cumulative rows (a
60%-of-dates cut left 143 fit rows under the fitter's 200 minimum; still
strictly temporal):
hits n=1140 bias +0.0868 brier 0.2626 -> 0.2511 d -0.0115 CI [-0.0139,-0.0097]
TB n=1050 bias +0.0834 brier 0.2490 -> 0.2438 d -0.0052 CI [-0.0061,-0.0045]
rbi n= 630 bias +0.0164 brier 0.2011 -> 0.1965 d -0.0046 CI [-0.0092,-0.0010]
runs n= 597 bias +0.0410 no map fittable (173 fit rows < 200)
ALL FOUR REFUSE: 2-4 eval date-clusters against a floor of 40. The floor
is the order's own and was not relaxed to force a pass.
A NULL THAT SCORED ITSELF: the first run reported hits at Brier 0.5567,
worse than predicting 0.5 for everything. fitIsotonic returns null below
its minimum, applyIsotonic then returns null per row, and (null-1)**2 is
1 while (null-0)**2 is 0 -- so the "Brier" was silently just the win rate
(0.5684). This project's signature Number(null)===0 breach, in my own
measurement code. Now a hard refuse.
PHASE 4: the bias is NOT a uniform shift. Identical favourite-longshot
shape on all four stats -- near zero or negative at 0.5-0.6, rising to
+0.21 to +0.28 above 0.9. The counter is over-confident specifically
about its favourites, which is the population a user acts on. Gradient is
hits ~ TB > runs > rbi, not the TB > RBI > runs anticipated.
PHASE 5/6 NOT RUN -- both gated on a Phase 3 deploy that did not open.
PHASE 7, refusal accuracy, first real measurement: refused props are
FURTHER from a coin flip than graded ones (TB refusals went over 21.6% of
the time). The obvious explanation, that refusals concentrate on players
who barely played, was tested and does not hold -- refused mean 3.20 AB
vs graded 3.39, 6.6% vs 6.2% with <=1 AB. So we pass on what we have no
INPUT for, not on what we cannot call. Refusing to invent a number
without a reference stays correct; the pass is not landing on the
genuinely uncertain props.
p_win never mutated, no p_win_calibrated written since nothing deployed,
no Bonferroni slot consumed. Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Nothing proved. For RBI even the ARCHETYPE split is theatre, so the honest
grade is the POOLED base rate.
PREMISE NOTE: the order's closing line says the batter board is
per-archetype-graded after this. Nothing has been rescaled for hits or
total_bases either -- no archetype slot has ever reached sample and
gradeBands remains built, gated and unwired. This is the fourth stat
measured, not the completion of three.
AUDIT: RBI 935 clean / 43 games; RUNS 617 clean / 33 games. Zero
quarantined. No archetype slot reaches 500 -- and the signature
archetypes the order names are the two SMALLEST slots on the board,
RBI->DRIVER at n=24 and runs->CATALYST at n=9. RUNS is refused
structurally before any factor is tested: 33 game clusters against a 40
floor.
INPUTS RECONSTRUCTED rather than declared missing. lineup_context only
covers 08-04 onward while settled rows start 07-31, so 187/617 runs rows
joined. But the play-by-play cache runs from 05-01 and the batting order
IS the order batters first appear -- slot, power-behind and reach-base
all rebuilt point-in-time, coverage 187 -> 574.
RBI, all THEATER: risp_opportunity +0.0047, extra_base_skill +0.0010,
risp x extra_base +0.0056. RUNS, all refused on clusters and all pointing
the wrong way: +0.0043 / +0.0008 / +0.0054.
THE COMPOUND IS THE WORST VERSION IN BOTH STATS. The causally-correct
compound was the most promising factor on the sheet and is the most
harmful in each. Two multipliers that individually carry nothing do not
cancel -- they compound each other's noise. Distinct from the
collapsed-sequence lesson: there the product of two REAL effects was too
small to use; here the product of two NULL effects is worse than either.
THE ARCHETYPE DOES NOT RESCUE IT, and this is where the session nearly
went wrong. The base rates look strongly differentiated (RBI DRIVER 0.609
vs BOMBER 0.413; runs GHOST 0.716 vs BOMBER 0.460). Gated directly
against the pooled base rate: RBI +0.0010 CI [-0.0034,+0.0050] THEATER;
runs -0.0028 CI [-0.0147,+0.0108] candidate at k=33. DRIVER's 0.609 is
n=23 -- small-slot noise wearing a decimal point. Read off the table
instead of gated, this would have shipped as "archetype differentiation
is real and large". It is not.
THE CROSS-STAT PATTERN THAT IS REAL -- the counter over-predicts every
batter counting stat measured:
total_bases p_win 0.5698 vs actual 0.5074 bias +0.0624
rbi p_win 0.4860 vs actual 0.4313 bias +0.0547
runs p_win 0.5949 vs actual 0.5749 bias +0.0200
Across four stats and three sessions, calibration is the systematic
defect and factor scarcity is not. TB's held-out isotonic fix (-0.0039)
still outperforms every factor tried on any stat, all null or theatre.
NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample, and for RBI the archetype split is itself theatre -- so the
honest band is the pooled base rate, which gradeBands returns by
construction.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PREMISE CORRECTION: the per-archetype rescale is not "proven and live on
hits". gradeBands was built, gated and explicitly NOT wired two orders
ago -- no hits archetype slot reached sample, every band came back
base-rate, and only defense_by_direction proved pooled. This applies an
unvalidated-at-archetype-level method to a second stat.
FULL-HISTORY AUDIT: 988 clean settled TB rows (101 quarantined, 948 with
p_win), 341 players, and only 9 DISTINCT GAME DATES. No archetype slot
reaches 500 -- BOMBER 340, GHOST 147, BRUSH 55. Confirmed short on full
history, not a windowed artifact. The 9-date figure matters more than the
row count: ~49 games means any game- or venue-borne factor has almost no
replication here.
THE BASELINE HAD TO CHANGE, to a harder null. TB lines vary (1.5 on 559
rows, 0.5 on 345), so a per-line personal base rate would rest on ~2 rows
per player-line and would have to be invented. The null is the counter's
own p_win, which already prices the line -- beating the champion, not
beating "he's due".
THE UNITS BUG, caught, and it had produced the best result in the
programme. The first run reported barrel_rate at Brier -0.0095, the
largest improvement ever measured here. fromStatcastRow returns
barrel_pct as a FRACTION (0.06) while the raw table stores 0-100, so
(0.06 - 7.8) * 0.018 clamped EVERY row to the maximum negative shift.
That uniform downward push "improved" Brier purely by leaning on the
counter's over-prediction and contained no barrel information at all.
Same family as the S80 trap, inverted. exit_velo was a second bug -- the
column is avg_exit_velo, so it read null on every row and reported n=0. A
zero is a wiring bug until proven an honest absence.
GATE with units fixed, 138 cumulative tests:
barrel_rate n=707 shift 0.0364 brier +0.0036 THEATER
exit_velo n=707 shift 0.0229 brier +0.0022 THEATER
hard_contact_allowed n=707 shift 0.0260 brier +0.0033 THEATER
park_weather_hit_type n=651 36 entities PENDING (k<40)
platoon_severity n=481 PENDING (n<500)
THE PREDICTED INVERSION WENT THE OTHER WAY. BOMBER x barrel_rate is
+0.0114, the single most harmful cell in the table, exactly where the
strongest proof was predicted. GHOST +0.0012. All sample-blocked so not a
verdict, but recorded so it is not claimed later.
AND IT IS NOT DOUBLE-COUNTING -- tested and refuted: corr(barrel, p_win)
= -0.061, the counter is not pricing barrel at all. The duller answer is
corr(barrel, counter RESIDUAL) = -0.012. Barrel is a real skill that
carries no information about what the counter gets wrong at this line.
That also closes the S81 lead: hard_hit r=0.153 at n=295 drifted to 0.135
at n=383 and is THEATER at n=707.
THE REAL FINDING: TB is miscalibrated, not under-factored. mean p_win
0.5698 vs actual 0.5074, bias +0.0624. Held out on a strict time split
(fit < 2026-08-02, eval 651 unseen rows): raw 0.25007, constant de-bias
0.24740 (-0.00267), isotonic 0.24621 (-0.00386). Worth more than any
factor tested and the only intervention pointing the right way -- and
still refused at the corrected bar on 32 clusters. A CANDIDATE, not a
result. It also explains the units bug's fake success exactly: a blanket
downward shift is a crude de-bias.
NO RESCALE. Nothing proved, nothing certified calibrated, no slot at
sample -- every band would be the honest base-rate band gradeBands
already returns by construction.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Nothing in this failed, which is what makes it the most instructive
negative so far. Link 1 proved (MAE 3.22 -> 2.80 batters faced). Link 2's
quality grain proved (2.70pp of realized separation). Both point-in-time,
both past cumulative correction. Their product is 0.37pp and detecting it
would take 52 seasons.
FRAMING CORRECTION: the order says Link 2 proved you can't predict the
reliever. Half true -- the INDIVIDUAL grain failed at 17.2%, but the
QUALITY grain PROVED. Pen-season-quality is a measured predictor here, not
a fallback after a failure.
TWO OF THREE SPECIFIED INPUTS COULD NOT BE USED HONESTLY. Pen archetype
did not prove (0.5669 vs a 0.5309 modal baseline, interval spanning zero)
so building it in would chain on an unproven link. And hitter
approach-identity -- "fastball-hunter", "finesse-vulnerable" -- does not
exist in this registry; MLB batter archetypes are BOMBER/GHOST/TORCH/
BRUSH/DRIVER/FLEX/ALPHA/HYBRID/CATALYST. Inventing one to condition on is
the fabrication the gate exists to catch. A power/contact split derived
from the sequence data was tested as a SEPARATE gated addition instead;
neither half proved.
GATE on the concentrated subset, 114 cumulative tests:
early-exit x WEAK pen n=1931 brier -0.0001 CI [-0.0014,+0.0010] NOT_PROVEN
early-exit x STRONG pen n=2574 brier 0.0000 CI [-0.0011,+0.0010] THEATER
all early-exit later ABs n=6869 brier -0.0001 CI [-0.0007,+0.0005] NOT_PROVEN
pooled all later ABs n=17891 brier 0.0000 CI [-0.0004,+0.0003] THEATER
Not pooled-diluted -- the concentrated subset was gated alone and is no
better.
THE CEILING, which explains it. The descriptive pass found the predicted
direction (+0.74pp weak pen, -0.79pp strong pen). The magnitude is the
problem and it is structural:
P(faces pen | early-exit flagged) 0.8075
P(faces pen | starter goes deep) 0.7149
exposure the flag actually buys 0.0925
hit-rate swing across pen quality 0.0394
MAX JUSTIFIABLE ADJUSTMENT 0.00365
actually applied 0.01930 -> 5.3x over-movement
A hitter's 3rd/4th plate appearance is ALREADY against the bullpen 71% of
the time when the starter is projected to go deep. Link 1 lifts it to 81%
-- nine points of extra exposure, not a change of opponent. The 5.3x
over-movement is precisely why the mirror subset reads THEATER rather than
as a small true effect.
A correctly-scaled version is not detectable either: 0.37pp is 0.37 SE at
n=1,931; the corrected bar needs n=168,488, an 87x shortfall, ~52 seasons.
STRUCTURALLY CLOSED, not sample-blocked. Waiting does not fix it.
NOT WIRED, and the self-check deliberately not wired either -- flagging
line-divergence on an adjustment measured as absent would advertise an
edge we just showed does not exist, which is fabricated reasoning one
layer up.
THE LESSON: link-by-link validation guarantees each link is real. It does
not guarantee the chain transmits anything. Size the multiplicative
structure BEFORE building -- one exposure term of 0.09 reduces a genuine
3.94pp signal to noise and no downstream care recovers it.
Link 3 confirmed skipped. Parallel track logged unchanged: TB n=948
pooled, BOMBER x TB 340, short by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The refinement was right. Naming the individual reliever failed; the same
question at the grain the chain needs passes, and it transmits more than
anything else measured in this chain.
WHY IT WAS WORTH RE-ASKING: last session's null (the pen is on average no
softer, +0.0010 on 35,760 PAs) does NOT rule this out, and treating it as
though it did would have been the error. An average washing out is fully
consistent with quality VARIATION mattering. It does -- actual arm quality
moves the hit rate monotonically across quartiles, 0.2244 / 0.2293 /
0.2410 / 0.2501, a 2.57pp spread, larger than the whole times-through-
the-order effect.
CLUSTER UNIT CORRECTED, THEN CHECKED RATHER THAN ARGUED. Last session
refused Link 2 partly as team-borne (30 bullpens, the park ceiling). My
first re-check was that 76% of pen-quality variance is within-team -- but
that is a statement about TREATMENT variance, not about where errors
correlate, and stopping there would have been picking the convenient
answer. Measured the actual thing: ICC of prediction error by team =
0.0261, design effect 1.41, SEs inflated ~19%. So the verdict was run
three ways:
unclustered CI [-0.0067,-0.0010] excludes zero
team-clustered (30) CI [-0.0086,-0.0003] excludes zero (below the
40-cluster floor -- indicative, not a pass)
design-effect adjusted CI [-0.0072,-0.0005] excludes zero
QUALITY GRAIN PROVES on the concentrated elevated-early-exit subset:
n=501 team-games, 426 clusters, MAE 0.0294 -> 0.0260, delta -0.0034, CI
[-0.0063,-0.0005] at 110 cumulative tests. Pooled also proves, so it is
not a subset artefact.
ARCHETYPE GRAIN DOES NOT: 0.5669 vs a 0.5309 modal-guess baseline,
corrected interval [-0.1073,+0.0268] spans zero. Two grains tested, one
earned a place -- penQuality.js exposes no archetype and a test asserts
it.
WHAT LINK 3 RECEIVES, which is the number that actually matters -- not
the MAE gain but realized outcome separation, prediction strictly
point-in-time:
predicted BEST pen 167 games 2,044 PAs hit rate 0.2231 +/-0.0180
predicted WORST pen 167 games 1,799 PAs hit rate 0.2501 +/-0.0200
2.70pp separated, intervals non-overlapping, capturing nearly all the
2.57pp available at the quartile grain. Caveat stated not buried: the
tercile cut is chosen in-sample; the prediction driving it is not.
BUILT: penQuality.js + 9 tests. Abstains below 5 prior club games and 40
arm appearances -- a league-average stand-in would assert "this is an
ordinary bullpen", which is a claim, and usually the wrong one for exactly
the clubs whose pens just turned over.
Link 3 is unblocked on a proven Link 2 at the quality grain only. Not run
here; this order scopes to building and gating Link 2.
Parallel track logged unchanged: TB n=948 pooled, BOMBER x TB 340, short
by 160.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The causal insight is right -- the game is a sequence and the matchup does
shift mid-game. The direction is backwards, measured on 93,663 plate
appearances from 1,238 games pulled free from statsapi.
LINK 1 PROVES. Starter batters-faced, point-in-time from his own prior
starts only, clustered on the pitcher: MAE 3.2226 -> 2.7990, delta
-0.4236, CI [-0.6006,-0.2731] at 0.9995 corrected for 107 tests, 1,706
starts across 204 pitchers. It finds the tail the chain needed -- early
exits are a 23.2% base rate, model-flagged starts are 34.0% early, lift
+10.8pp.
Scope correction inside Link 1: the order specifies fatigue x GAME
SCRIPT, but game script is not available at grade time -- whether he gets
hit tonight is the thing being projected, not an input to it. Only the
workload half is measured; the in-game half is recorded as a live feature,
out of scope, rather than quietly folded in.
LINK 2 DOES NOT PROVE, twice over. Model accuracy 17.2% vs an 8.6%
baseline -- doubling it sounds good and is not, since naming a specific
arm is wrong five times in six. And structurally the entity is the
BULLPEN: 39,629 post-starter plate appearances across 30 clubs is 30
readings, below the 40-cluster floor, the same permanent ceiling as park
geometry and team defence. LINK 3 NOT RUN, per the order's own rule.
THE PREMISE IS REFUTED, and this chains on nothing so it was safe to
measure:
vs STARTER n=48,492 hit rate 0.2444 +/-0.0038
vs BULLPEN n=35,760 hit rate 0.2373 +/-0.0044
The pen is 0.7pp HARDER. The specific effect the chain exists to exploit
-- early exit making later at-bats softer -- is +0.0010 on 35,760 PAs. A
well-powered null, not a sample problem.
What IS real is times through the order: TTO1 0.2351 -> TTO2 0.2515 ->
TTO3 0.2518. A starter does decay as the lineup sees him again, but that
advantage is SURRENDERED when he leaves, not extended -- the pen is
harder than his second and third time through. A modern bullpen is a
queue of fresh specialists throwing one inning each; there is no tiring
arm to punish.
So the insight survives inverted, and Link 1 stays valuable for the
opposite reason it was built: a likely early hook predicts the hitter
LOSES his third-time-through look (0.2518 -> 0.2373 on that PA). The
mispricing is on hitters who get an EXTRA look at a starter going deep.
BUILT: predictionGate.js + tests -- the two-part gate for a continuous
prediction. factorGate binarises outcomes for Brier, which would destroy
a target like batters faced. Same discipline, same THEATER verdict, real
scale.
PRE-REGISTERED NOT RUN: Link 2' using a PA-weighted bullpen AGGREGATE
rather than a named arm. Recorded rather than substituted in -- running
Link 3 on a swapped-in Link 2 is the assumed-link failure the order
forbids. Given the premise result its expected value is now low.
PARALLEL TRACK logged: total_bases n=948 pooled, BOMBER x TB 340, short
by 160. Sample-readiness only, not a verdict.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The premise does not hold. proven-status.js run fresh: PROVEN_SET is
EMPTY, no archetype x stat reaches the gate. pitcher_contact_profile has
a CI upper bound of exactly 0.0000 and platoon_severity is held on
4.5%-contaminated splits, so the proven set is one factor, pooled, not
three archetype-conditioned ones. The specific pattern the order names --
defense strong for GHOST/BRUSH, null for BOMBER -- is the one I measured
running the OTHER WAY yesterday, both noise-dominated.
But the second blocker is new and matters more, because it would stop the
rescale even if the factors had proved: the grade does not separate
within any archetype. Every archetype collapses to ONE band at the
corrected bar, because bands merge when their intervals overlap and
publishing two letters we cannot tell apart is a distinction we have not
measured.
Uncorrected, so the ranking is visible rather than hidden by the bar,
this INVERTS the order's design. The order gives contact types the
factor-rich treatment and power types honest base-rate, reasoning that
single-game hits are variance for a power profile. Measured:
BOMBER n=466 corr(p_win,outcome) +0.207 quintiles 0.75 0.62 0.60 0.48 0.48
GHOST n=192 corr(p_win,outcome) -0.007 quintiles 0.47 0.63 0.74 0.58 0.45
BOMBER is the one archetype the model ranks, and it splits into a real
A 0.660 / B 0.481 at 95%. GHOST is flat, and non-monotone -- its most
confident reads hit 47% while its middle reads hit 74%. Shipping as
specified would have given the factor-rich treatment to the archetype the
model reads worst and left base-rate on the one it reads best. That is
mechanically sensible in hindsight: a power hitter's hit tracks whether
he can damage the arm, a contact hitter's depends on balls finding holes.
BOMBER's split does not survive the cumulative correction at 106 tests.
Exposing it by loosening the correction is the curve-to-make-A's the
order forbids, so it stays one band.
BUILT: gradeBands.js -- lift against the archetype's OWN base rate (the
same 62% is lift for a 45% profile and a deficit for a 68% one),
indistinguishable neighbours merged, thin bands PROVISIONAL not dropped,
Wilson intervals widened by the cumulative correction. The two-bar rule
is structural: proven-alone, calibrated-alone and neither all return
base_rate with the reason stated, so with nothing proven no
factor-informed band can be produced at all.
reasoning() is built and tested but NOT wired to the card -- there is no
per-archetype band being served, so attaching the copy now would ship
product language for a rescale that does not exist.
NOT BUILT: the specified power-type reason "the matchup edge is in
total_bases". total_bases is recorded INCONCLUSIVE (+0.0038, CI
[-0.068,+0.075]). Wiring it would assert an edge measured as
indistinguishable from zero -- the exact fabricated-reason failure this
module exists to prevent.
BOMBER x hits is 29 rows short of the gate and is the archetype the model
actually reads. That is the first slot to test, not GHOST.
Counter and frozen clusters byte-identical. No letter was moved.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
decided everything
The premise does not hold. prove-hit-factors.js has no date filter
anywhere in it and pages the full table -- there was never a window to
widen. Full clean history is 1,266 rows, not 2,715. platoon was not
"proved" last session, it was explicitly held on 4.5%-median-contaminated
season-to-date splits, and pitcher_contact_profile was demoted. The
proven set going in was one factor, not three.
STEP 1: no archetype slot reaches n>=500 on full history. Best is BOMBER
at 408, and BOMBER is the most common archetype on the board. GHOST 173,
BRUSH 64, DRIVER 43, CATALYST 16. These are confirmed genuinely short,
not artifacts.
STEP 2 is where the real finding is. park_hits initially PROVED at 619
rows across 45 games -- but those games only ever visited 14 distinct
park values. A park effect is replicated across parks, and unmodelled
park heterogeneity is confounded with the thing being estimated. Each
factor is now clustered on the coarser of the game and the entity its
treatment rides on.
That flipped two verdicts and confirms Kev's causal-correctness thesis
from a new direction: defense_by_direction has 442 hitter-team units of
replication where crude team defense has 26. The correct atom is not just
more accurate, it is the only one measurable at all. park_hits (14) and
defense (26) can never be validated however long the ledger runs -- the
same ceiling as park dimensions, reached independently.
Also fixed a bar I got wrong last session: I transplanted the 500-row
floor onto clusters, which refused a factor with 1,059 rows over 85 games
while answering neither question. Two floors now -- rows>=500 for a stable
estimate, clusters>=40 for a trustworthy interval. Not a lowered bar:
park_hits and defense are still refused.
PROVEN: defense_by_direction only, pooled, [-0.0054,-0.0012] at 99 tests.
It stays POOLED-ONLY -- no per-archetype reasoning wired, nothing
grandfathered. The card must not say "GHOST: defence matchup strong"
because we have not earned that sentence. The predicted fingerprint did
not appear either: BOMBER -0.0036 vs GHOST -0.0024, the opposite
direction, both noise-dominated. Recorded so it is not claimed later.
RESCALE: NOT READY. One proven factor worth -0.0031 Brier. Rescaling on
that is relabelling.
Counter and frozen clusters byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The platoon test's n=452 described how much of the JOIN survived, not how
much data exists. There are 1,266 clean settled hits rows and zero
quarantined ones. platoon_splits had been ingested from tonight's lineups
only (315 players), so any hitter who settled a prop without appearing in
an ingest-day lineup was silently absent from every test.
Backfilled all 380 hitters (81 fetched, 0 unresolved). Re-ran on 1,059
rows, up from 452.
THE DEMOTION IS THE HEADLINE. pitcher_contact_profile, the strongest
proven factor in the programme (-0.0064, CI [-0.0113,-0.0014]), roughly
halved to -0.0034 on more than double the sample and its corrected
interval now spans zero. The Bonferroni denominator also rose to 55,
which widens every interval -- but a denominator cannot move a point
estimate, and that halved on its own.
platoon and platoon_severity now clear the bar and are NOT promoted.
Upper bound -0.0001, on season-to-date splits that contain the games they
predict: measured contamination is 4.5% median, 12.4% at p90, 137% worst.
I had assumed ~1%. They stay CANDIDATE pending point-in-time splits.
GAME-LEVEL IS A DIFFERENT PROBLEM. game_context held zero weather rows
ever -- not because the fetcher was wrong (it correctly targets
Open-Meteo's archive) but because ledger_entries keys a game as
mlb:2026-08-03:Away@Home and game_context keys it as mlb:823437. Every
lookup missed and NULL columns read as honest absence. Third occurrence
of that class.
Fixed the join: 96/101 settled games now carry actual archived weather,
park dimensions backfilled 15 -> 30 venues.
But 928 total_bases rows sit on 47 games at 17.6 rows per game. Park and
weather assign one value per game, so resampling rows would have
manufactured a pass. factorGate now resamples clusters when rows carry
one and judges sample against effective_n; unclustered rows keep the
original path byte-for-byte. Verdict: 47 clusters < 500, and the point
estimate is +0.0011 -- worse, not merely unproven.
Weather needs ~57 more days. Park dimensions need never: there are 30
ballparks in MLB, so a venue-constant factor can never reach 500
independent units. That bar was built for player-level factors and does
not transfer.
Wind is refused. We have speed and bearing for all 96 games; we lack park
orientation, and 220 degrees is blowing out at one park and in at
another. Using speed alone would assert an effect while discarding the
sign that decides what it is.
Counter and frozen clusters untouched.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Two fixes in the weather path, and the second was hiding behind the first. The
scalar weather_mod cannot express a hit-TYPE conversion at all -- wind out and
warm turning fly balls into extra bases, and cold heavy air turning them into
outs, collapse to the same number once multiplied -- so the raw temperature,
wind speed and wind direction are now retained alongside it.
And the old guard only kept the environment when the multiplier was not 1,
which silently discarded the forecast for every ordinary night. That is the
majority of games, and precisely the rows a hit-type model would need in order
to learn what ordinary looks like.
Platoon severity is built and measured at n=452, which is 48 rows short of the
gate: CANDIDATE_PENDING, neither proven nor theatre. It moves less than flat
platoon (0.021 against 0.026), consistent with the pattern, and its Brier point
estimate is favourable but the corrected interval still spans zero.
Worth naming: the refusal costs sample, and that is the design working. Flat
platoon scores 741 rows because it will happily apply a boost to anyone;
severity scores 452 because the other 289 are hitters whose split we cannot
actually read at 60 plate appearances on the short side. Buying those rows back
by shrinking instead of refusing would have produced a number indistinguishable
from a measured league-average split, which is a different claim from the one
the data supports.
Park dimensions are ingested and verified in production across fifteen venues,
joined by the venue the game is actually at rather than inferred from the home
team -- neutral-site and international games break that assumption without
surfacing an error.
The park-and-weather-to-hit-type atom is NOT built. Its inputs landed this
session and carry a single as_of date, so testing it on total_bases would be
scoring games with inputs that postdate them. Building it now would produce
something plausible rather than something proven.
Proven factors for hits remain pitcher_contact_profile and
defense_by_direction. 4,307 tests green (344 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Applying the method that worked for defence to the two factors the code flagged
as still crude.
PLATOON. The flat version is 'lefty versus righty, add a boost', and it failed
the two-part gate for the same reason team-average defence did: it is not the
unit the causal story runs through. The advantage is only worth what THIS
hitter's split is actually worth -- measured on a real hitter, .284 against
left-handed pitching versus .221 against right-handed, a 63-point split, where
the flat factor applied the same six percent to him and to a hitter with none.
Most of the work is sample discipline, and the second rule matters more than
the first. Severity shrinks toward the league split weighted by the SMALLER
side's plate appearances, because a 500-against-40 split is a 40-PA read. And
below a floor it REFUSES outright rather than shrinking, because a
heavily-shrunk severity is indistinguishable from a measured league-average one
and those are different claims -- without the refusal the atom would quietly
assert a league-typical split about every September call-up in the league.
Switch hitters turn out to be the easy case misread as the hard one. He bats
opposite by choice so the direction is never in doubt, but the per-side value of
his swing is a different question and one this sample cannot answer, so he is
unreadable rather than credited with an automatic edge.
PARK DIMENSIONS. Free from statsapi's venue endpoint, which carries fence
distances, roof, turf and elevation outright -- Wrigley returns 355 down the
left line, 400 to centre, 353 to right, at 595 feet. parkFactors holds run
COEFFICIENTS, which structurally cannot express a park that turns outs into hits
without scoring, and that is why the crude park factor failed.
The park join is by the venue the game is ACTUALLY at, carried from the schedule
feed, never inferred from the home team -- neutral-site and international games
break that assumption and they break it silently. A venue with no geometry at
all is absent rather than a park with zero dimensions.
Both tables dated in the primary key. Venue geometry changes rarely but it does
change, and by now that is the default rather than a lesson.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Kev's insight holds, and the data says so cleanly. defense_by_direction PROVES
on hits -- n=528, Brier -0.0034, interval [-0.0059, -0.0009] at the 99.9% level
the cumulative correction now demands -- while team-average defence remains not
proven, its interval still spanning zero. Same signal, same rows, different
unit.
The detail worth keeping is that the causally-correct atom moves the number
LESS THAN HALF as much as the crude one, 0.013 against 0.030, and is the one
that is reliably right. The team average was moving more and knowing less. Big
movement is not evidence of a good factor; it is frequently the tell.
Both halves turned out to be free, as the order expected. Savant's batted-ball
leaderboard carries pull/straight/oppo crossed with ground/air for 609 hitters
-- the statcast leaderboard we already pull does not, it has nineteen columns
and no direction at all -- and the OAA feed already carries each fielder's
position, so per-position defence is a regrouping of last week's ingest rather
than a new source. Verified in production: 609 spray profiles, 31 teams.
Handedness is what joins them and getting it backwards would have been
invisible. Pull for a right-handed hitter is the left side; for a left-handed
hitter it is the right side. A model that ignored `bats` would send half the
league's grounders to the wrong infielders and still look like it was reading
defence, and nothing downstream would have caught it. Switch hitters bat
opposite the pitcher, which this does not resolve, so they are unreadable
rather than guessed.
Unmeasured zones are renormalised away rather than contributing a zero, since a
zero asserts an exactly-average fielder standing there, and coverage states
honestly what share of a hitter's contact we could actually read.
ATOM 2 is input-blocked rather than sample-blocked, and the distinction matters
because waiting will not fix it. The weather free-source check passes --
Open-Meteo is already wired and exposes temperature, wind speed, wind direction
and precipitation -- but those raw fields are collapsed into a single scalar
modifier and wx_forecast is empty on all 1,119 settled rows. Park DIMENSIONS
are not ingested at all; parkFactors holds coefficients, not wall heights or
fence distances. A park-and-weather-to-hit-type conversion needs both, so it is
scoped rather than half-built: retaining the raw weather fields is the cheap
half, dimensions are the missing one.
Proven factors for hits are now pitcher_contact_profile and
defense_by_direction, both pooled; every per-archetype slot remains
sample-blocked.
4,297 tests green (342 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Team-average defence failed the two-part gate for hits, and the reason was the
unit rather than the signal. A left-handed pull-ground hitter meets the first
baseman and the second baseman and almost nobody else, so a team total averages
in five fielders who will never touch his ball.
Both halves were already free on the host we pull from. Statcast publishes
spray x trajectory per hitter -- pull/straight/oppo crossed with ground/air,
608 hitters -- and the OAA feed already carries each fielder's position, so
per-position defence is a regrouping of data ingested last week rather than a
new source. Zero new sourcing, as the order expected.
Handedness is what joins them and getting it backwards would be invisible: pull
for a right-handed hitter is the left side, pull for a left-handed hitter is the
right side, so a model ignoring bats would send half the league's grounders to
the wrong infielders and still look like it was reading defence. A switch hitter
bats opposite the pitcher, which this does not resolve, so he is unreadable
rather than guessed.
Two properties the crude version could not express, both locked by test: two
teams with the SAME total defence read differently for a pull hitter, and a
ground-ball hitter and an air hitter read the same team in opposite directions.
Unmeasured zones are renormalised away rather than contributing a zero, which
would assert an exactly-average fielder standing there, and states
honestly what share of a hitter's contact we could actually read. Nothing
readable at all returns null, so the caller falls back to the base rate instead
of to an invented 1.0 that looks measured.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The question was whether the hit grade reads tonight's game or just says he is
due. Answering it needed a gate that correlation cannot provide, because
correlation cannot separate the two ways a factor looks alive: it reads the
game, or it moves the number and reads nothing. The second is what a product
ships by accident -- arch-v1 moved 76% of rows by 2.5 points, changed
resolution by 0.0000, and was live for months, and no user could have told.
So a factor must now clear both conditions: move the prediction off the
player's own leave-one-out base rate, AND improve out-of-sample Brier. Brier
rather than correlation, because correlation asks whether the ordering improved
and this asks whether the NUMBER got closer to what happened -- and for a graded
probability the number is the product.
The correction applies to the interval itself, which turned out to matter more
than expected. A plain 95% CI is the right bar for one test; at fifty
cumulative tests roughly two or three intervals exclude zero by chance alone.
Widening to 1 - 0.05/tests, currently 99.9%, flipped both defence and platoon
out of "proves". A 95% interval would have shipped two unproven factors into
the grade, with reasoning text explaining them to users.
That forced a distinction I had initially collapsed. Defence and platoon have
FAVOURABLE point estimates whose corrected intervals merely span zero, and
calling that THEATER would repeat the error this codebase keeps correcting:
insufficient evidence is not evidence of absence. THEATER is now reserved for
its one real meaning -- moves the number, reads nothing -- and
NOT_PROVEN_AT_CORRECTED_BAR names a real candidate held to a bar that rises with
every hypothesis the programme tests.
Result on 741 settled hits rows: pitcher_contact_profile PROVES, improving
Brier by 0.0066 with a 99.9% interval of [-0.0114, -0.0016]. Defence (-0.0043)
and platoon (-0.0039) are not proven at the corrected bar. Park is
sample-blocked at n=405. Zero factors are theatre, which is the genuinely good
news: nothing decorative is being wired. Per-archetype every slot is
sample-blocked (BOMBER 252-294, GHOST 67-125).
Two spec gaps worth recording. The approach identities the order names -- SPRAY,
DAMAGE-DEALER, COUNT-WORKER -- do not exist in the registry; the MLB batter
archetypes are BOMBER, GHOST, TORCH, BRUSH, DRIVER, FLEX, ALPHA, HYBRID and
CATALYST. And parkFactors maps hits to run_base, so there is no hits-specific
park factor at all: a park that turns outs into hits without producing runs is
invisible to the input we have.
The grade rescale is NOT run. It was explicitly gated on the factor proving,
and one pooled factor worth 0.0066 of Brier is not a factor-informed
distribution -- rescaling on it would dress a base-rate model as a matchup
model, which is the exact thing this gate was built to prevent.
4,286 tests green (340 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Fitted the isotonic map on game_date < 2026-08-02 (n=589) and evaluated it on
everything from that date forward (n=383). The map never saw the evaluation
rows, which is the only thing that makes the result mean anything -- fitting
and evaluating on the same rows always looks perfectly calibrated, because the
map is reciting the answers it was built from.
It works, on most of the distribution. Held-out after correction: 0.477 comes
back 0.506, 0.587 comes back 0.580, 0.667 comes back 0.603 -- against raw
errors of +0.191, +0.279 and +0.246 in the same bins. Ordering survived, and
that was verified pairwise rather than assumed, because a broken map would
silently destroy the one thing this model does well.
Two findings matter more than the pass.
First, the honest ceiling is 0.667. Once the numbers are truthful this model
has no 80%-plus hit reads at all -- the top of its range was miscalibration,
not confidence. A four-leg ticket at the ceiling is 0.198, where the raw
numbers implied 0.686. The high-floor parlay is a two-thirds-per-leg
proposition, and that is the number to say out loud.
Second, calibration is certified BY BAND rather than by a blanket flag.
Held-out error was -0.029 and +0.007 through the middle but -0.167 at the
bottom and +0.063 at the top: the model is trustworthy over most of its mass
and untrustworthy at both edges. A single true/false would either throw away
the 72% that works or ship the edges that do not. Only a probability inside a
certified band is marked stackable, and that flag is what chainAcross requires
before it will compound anything. The certified band is 0.40 to 0.60, n=276.
A methodological catch on the way: my first pass condition demanded honest bins
at 0.70 and above -- but honest calibration REMOVES those bins, since the
ceiling drops to 0.667. The gate would have failed the repair for succeeding.
It now tests the highest remaining band instead of a fixed threshold.
Wired forward with the same discipline: calibrationService fits strictly before
today, splits by time rather than at random, and returns null on thin history
so that "no calibrator" means nothing is stackable rather than "trust the raw
numbers". p_win is never mutated -- the calibrated value rides beside it as
p_win_calibrated, because a calibration map is a correction to a forecast, not
a different forecast, and the counter stays byte-identical.
4,275 tests green (339 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The order's own prerequisite for the hit-parlay surface was to verify the hit
probability is calibrated. It is not, and the failure is exactly the shape that
destroys a parlay.
Measured on 972 settled hits props: the model is monotonically over-confident
at the top and flat above 0.70. Predicted 0.911 comes back 0.630. Predicted
0.844 comes back 0.630. Predicted 0.747 comes back 0.605. There is no
discrimination at all in the range a parlay is built from, and the error runs
in the flattering direction. Four "91%" legs are 0.686 by the model and 0.157
in fact -- a 4.4x overstatement that compounds with every leg added.
Single props survive a calibration error of that size. A parlay multiplies it.
So chainAcross REFUSES to compound atoms not marked calibrated, and refusing is
the feature rather than a limitation: a ticket built on these numbers would be
confidently wrong in the direction the user pays for.
calibration.js provides the reliability table, the gate (tolerance 0.05,
weighted to the high end because that is where tickets live) and an isotonic
fit. Isotonic is the honest repair here because it is monotone: the model's
ordering survives untouched while the numbers move to what actually happened.
The fitted map says 0.65 -> 0.594, 0.85 -> 0.639, 0.91 -> 0.639.
chain.js is the portable core -- base events plus context, through a chain
function, into a PLUGGABLE aggregator: across players for a compound ticket, up
to the team for expected scoring. The sport-specific parts are inputs rather
than code paths, so basketball plugs in as content. The archetype
redistribution hook is there now, dormant in baseball because a nine-run lead
does not change who bats next, and live in basketball where a blowout fades the
star and feeds the bench.
Two judgement calls worth naming. Treating same-game legs as independent errs
in the FLATTERING direction, since they share pitcher, park and weather -- so
correlation shifts the compound toward the weakest leg, bounded, and is labelled
an approximation rather than a joint distribution. And market divergence does
NOT downgrade confidence: it flags a contested script whose props are either the
best or the worst on the board, and which one is unknown until settled.
Internal inconsistency does downgrade it, because per-entity reads failing to
sum to the team read means one of them is wrong and we do not know which.
Not built: the independent game-script projection. It needs proven team-level
atoms and out-of-sample validation against actual margins, and no atom has
passed the gate yet. Building it now would produce something plausible rather
than something proven, which is the failure mode this whole programme exists to
avoid.
4,269 tests green (339 suites); web build exit 0; counter and frozen clusters
byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Verified in production via the new on-demand endpoint: 153 batting-order rows
across 10 games (orders 1-9), and 149 hitter-opportunity rows with RISP shares
ranging 0.170 to 0.528.
The data passes its own coherence check on arrival: the highest RISP-share
hitters all bat fourth and fifth, which is exactly where the mechanism says the
RBI opportunity lives. Nothing was fitted to produce that -- it is the two
tables joining and agreeing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The first prod run wrote zero rows while the parser demonstrably works locally
(144 rows, 10 games with lineups posted), so the zero was wiring rather than
absence -- but diagnosing that required a full snapshot, which now takes about
three minutes and 524s at the edge.
Same reasoning as the statcast refresh endpoint: a job is proven by running it
and reading the result, never by waiting for the slot it rides in. This makes
the ingest verifiable in seconds, so 'zero rows' can be told apart from 'no
lineups posted yet' immediately -- which is the exact confusion the defence
ingest hit when a doubled path 404'd and read as 'Statcast has no fielding
data'.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
RBI is power TIMES opportunity. The same swing drives in one run or three
depending on who is on base, and a hitter batting with the bases empty cannot
drive anyone in however hard he hits it. Every context-free model of RBI here
has failed, and the failure kept being read as 'skill inputs don't work for
RBI' when the truth was that we were modelling half the stat.
Both halves are free from statsapi.mlb.com, which we already call for game
logs, schedules and probable pitchers. No new provider, no key, no quota.
RUNG 1, batting order: schedule?hydrate=lineups returns homePlayers and
awayPlayers as ORDERED arrays of nine, and the order IS the batting order --
index 0 is the leadoff hitter. That single fact gives CATALYST its identity
and supplies lineup-position context for every context-dependent stat.
RUNG 2 turned out cheap, which the cheapest-first rule did not expect. It
looked like it would need play-by-play reconstruction across a season; statsapi
serves situational splits directly, so 'how often does this hitter bat with
runners to drive in' is ONE call per player rather than one per game. Measured
on a real hitter: 87 plate appearances with runners in scoring position
producing 25 RBI, against 302 with the bases empty producing 17. That ratio is
the opportunity half of the stat and it is the thing no amount of exit velocity
can tell you.
Both tables are dated in the primary key. statcast_aggregates was built
upsert-in-place and that silently made every backtest leak the games it was
predicting; a lineup is worse still, because it is a PRE-GAME fact that changes
by the hour, so an in-place table would overwrite what we knew at grade time
with what turned out to be true.
Absent stays absent throughout: no lineup posted is an empty slate rather than
a guessed order, a short lineup records fewer slots rather than padding to
nine, and a hitter with no splits is null rather than a zero RISP share --
which would assert he never bats with runners on, a strong claim and usually a
false one.
Wired into the snapshot best-effort, so a context failure can never break the
pipeline it rides in. The three pre-registered theories are now marked
input-ready rather than input-blocked: DRIVER's power x runners-on and power x
lineup-position, and CATALYST's speed x on-base x power-behind. They are
sample-blocked from here, and the proofs run under native cumulative
correction as sample accumulates -- ingesting is not proving.
Counter and frozen clusters byte-identical. 4,250 tests green (338 suites);
web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Every slot in this batch is far below the gate -- DRIVER x hits 39, CATALYST
under 22, SINKER not yet gradeable at all -- so per the order's own sample rule
these are CANDIDATE-pending-accumulation, not tested-and-failed. Testing them
now would produce noise and burn cumulative-correction budget on it.
What IS deliverable is the input SINKER's theory needs, and it turned out to be
free. The OAA feed already carries each fielder's position, so infield-only
defence is derivable from data ingested yesterday: 1B/2B/3B/SS summed
separately from the outfield. Team-total OAA is the wrong unit for a
ground-ball pitcher -- he lives on the infield converting grounders and his
outfielders are close to irrelevant to him, so averaging them in dilutes
exactly the signal. On a real team the split shows a +15 infield inside a +2
team total, which is the dilution made visible. Under three measured fielders
in a unit is absent rather than zero, same rule as everywhere else.
The three theories are now PRE-REGISTERED in the registry with their mechanism
and the skill each would validate, marked CANDIDATE. That is the point of
writing them down before the sample exists: the claim is on the record with a
date and cannot be quietly reshaped into whatever the numbers turn out to
support once they arrive.
Two of them are input-blocked rather than sample-blocked, and the distinction
matters because waiting will not fix them. DRIVER's RBI theory needs
baserunner state and CATALYST's runs theory needs both baserunner state and
batting order; we ingest neither, and player_role_profiles is empty. So RBI
does not unblock on DRIVER -- it unblocks on ingesting lineup context, which
is a sourcing question, not an accumulation one. SINKER is the only one of the
three whose inputs are now ready.
Premise note: no registry re-adjudication demoted anything last session. The
proven set was empty, zero features were demoted, nothing was recalibrated,
and nothing is published -- node scripts/proven-status.js confirms it in one
command.
Counter and frozen clusters byte-identical. 4,238 tests green (337 suites);
web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
There is nothing to re-adjudicate. The proven set is empty and always has
been -- verified three ways: proven-status reports EMPTY, validatedSkills()
returns {} for every archetype, and zero conditioning entries have ever
reached PROVEN. The one PROVEN feature is recent_frequency_prior, which is the
incumbent counter itself, proven by the S78 ablation as ~100% of the
champion's resolution. It is the baseline every challenger is measured
against, not a conditioning interaction, and demoting it would leave the model
with nothing to grade from.
A correction to the premise: the cumulative gate did NOT catch a false
positive last session. It caught nothing, because there was nothing in the
proven set to catch. What it did was tighten alpha from 0.0026 to 0.0013
within one session, which demonstrated the mechanism working rather than a
demotion. So steps 3 and 4 -- demote, recalibrate -- are vacuous here, and
readjudicateAll says so plainly rather than glossing a no-op.
But the worry behind the order was well founded, and the audit found the real
exposure: promote() did not require the cumulative denominator. It checked n,
lift and CI, and nothing stopped a future session from testing eight
hypotheses, correcting by eight, and promoting on a p-value that would not
survive the programme's real denominator. That is precisely the hole that
makes a retroactive re-adjudication pass necessary later, so it is closed at
promotion time instead. isSufficient now refuses evidence carrying no
correction, evidence corrected against fewer tests than the cumulative count,
and any p-value that does not clear 0.05 over its own test count. The same
rule guards a PROVEN conditioning entry.
The second audit found two of four analysis scripts still correcting
per-session; pitcher-prove-k and tb-solo-and-interactions now use the
cumulative ledger, so the correction is native on every path.
reAblation.js is the standing second line: pure and injectable, so the
decision rule cannot drift from the gate's, and every verdict records both
p-values and both test counts so a demotion is re-derivable by anyone. A
feature promoted at alpha 0.05/20 can demote on the same p-value once the bar
is 0.05/60 -- correct, because the bar rose only after the programme had more
chances to get lucky. No fresh measurement is PENDING_RETEST and never a
demotion: absence of a re-test is not evidence, and demoting on it would
punish whichever stat happens to be off-season.
Net effect on the proven set is zero. No demotions, no recalibrations, and no
public ledger event -- announcing "recalibrated after re-adjudication" when
nothing changed would itself be a false signal of rigour.
4,238 tests green (337 suites); web build exit 0; counter byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Two things shipped that stand regardless of sample.
DEFENCE. Statcast Outs Above Average is free on the host we already pull six
feeds from, so there was nothing to decide. 514 fielders, aggregated to team
level -- the unit a batter's prop actually needs, the defence behind the
pitcher he faces -- and persisted as 31 team rows. Verified in production.
Cubs +56 best, Mariners -29 worst.
Unknown is not zero, and it bites unusually hard here: an OAA of 0 is a REAL
reading meaning exactly average, so coercing absence to 0 would assert that
every unmeasured fielder is league-average, which is the commonest defensive
profile there is. team_defense also carries as_of_date in its primary key from
the first row -- statcast_aggregates was built upsert-in-place and that
silently made every backtest leak the games it predicted, so point-in-time is
available here before it is needed rather than after a wrong answer.
A bug worth recording as a class: BASE already ends in /leaderboard, so the
new feed built a doubled path and 404'd. Because a failing feed degrades to an
empty index by design -- correct, so one broken source cannot fail the whole
pull -- it surfaced as "fielding_oaa: 0 rows", which reads exactly like
"Statcast has no fielding data". Graceful degradation makes a wiring bug look
like an honest absence.
CUMULATIVE CORRECTION. Bonferroni had been applied per session throughout: a
run testing eight features corrected by eight. Across a programme's lifetime
that is wrong in the dangerous direction, because every order gets a fresh
generous alpha and the false-positive rate compounds quietly. Correcting by 8
when sixty have been tried is how a noise result eventually gets recorded as
PROVEN with a p-value to point at. The denominator is now distinct hypotheses
ever tested, persisted, and it moved 19 -> 38 within this session alone, alpha
0.0026 -> 0.0013. Re-tests deliberately do not inflate it: re-asking the same
question on more data is not a new shot on goal, and counting it would punish
the discipline of waiting for sample.
THE MEASUREMENT. The differential the theory predicted is present: defence
correlates with the counter's residual at +0.130 for GHOST, the contact and
speed archetype, and -0.018 for BOMBER, the power archetype. A GHOST's hits
depend on whether anyone can range to the ball; a BOMBER's barrels clear the
defence entirely. So a flat BOMBER result is the theory working rather than
the test failing.
It is not a result. GHOST is n=104 against a 500 bar, with p=0.188 against a
corrected alpha of 0.0013 -- three orders of magnitude short. Both are
recorded as CANDIDATE with their measured lift, tagged contact-skill, so the
re-run at full sample compares against a recorded baseline.
Nothing proved, so nothing was recalibrated and nothing shipped.
4,228 tests green (336 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
BASE already ends in /leaderboard, so the new feed built
.../leaderboard/leaderboard/outs_above_average and 404'd. Because a failing
feed degrades to an EMPTY index by design -- correct behaviour, so one broken
source cannot fail the whole mechanism pull -- it surfaced as 'fielding_oaa: 0
rows' rather than as an error, which reads exactly like 'Statcast has no
fielding data'. Worth noting as a class: graceful degradation makes a wiring
bug look like an honest absence.
Verified: 514 fielders, 31 teams. Cubs +56 OAA, Mariners -29.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Defence was the one conditioning category with no derivable proxy: nothing we
ingest measures fielding, and a team's pitchers' hits-allowed conflates
pitching with defence and would validate the wrong skill. Statcast publishes
Outs Above Average free on the same host as the six feeds already pulled --
verified live at 513 fielders -- so there was nothing to decide.
Added as a seventh feed, indexed per fielder and aggregated to team level,
which is the unit a batter's prop actually needs: the defence behind the
pitcher he faces. Summed OAA is the team's outs converted above average; the
mean rides along because a team with more measured fielders would otherwise
look better merely for being measured more, and a team with under three
measured fielders is absent rather than thin.
Unknown is not zero, and it bites unusually hard here: an OAA of 0 is a REAL
reading meaning exactly average, so coercing absence to 0 would assert that
every unmeasured fielder is league-average -- the most common defensive
profile there is, and a fabricated fact rather than a neutral default.
team_defense carries as_of_date in its primary key from the first row.
statcast_aggregates was built upsert-in-place with a single as_of date, which
silently made every backtest leak the games it was predicting and cost a
session to find; this makes point-in-time available before it is needed
instead of after a wrong answer.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The order opens with "two proven clusters live". They are not proven -- the
proven set is empty -- and this is the fourth consecutive order to start from
a stronger claim than the measurements support. Correcting that in prose four
times has not worked, so this session adds scripts/proven-status.js, which
recomputes the answer from the ledger: hits LOSES (-0.096, CI excluding zero),
total_bases INCONCLUSIVE (+0.004), strikeouts INCONCLUSIVE (+0.259 at n=57).
It deliberately reports sample readiness separately from recorded verdicts, so
"n>=500" can never again be read as "passed".
A counting error worth recording. The first read of the top-volume archetype
said BOMBER x hits was 641 rows -- gate-ready. It is 287. model_snapshots
holds one row per prop PER SNAPSHOT CYCLE, so joining it to ledger_entries
counts each ledger row once per cycle it appeared in. Deduping on the ledger
row id gives the true figure, and my own status script had the same bug until
it was fixed. That is the difference between running the gate and being short
by 213.
So no archetype x stat combination reaches the gate. BOMBER x hits at 287 is
the closest; pitcher archetypes are untestable at 58 settled strikeout rows
across all of them, so the pitcher half of this order could not be run.
The registry is built: recordConditioning keys archetype x underlying-skill x
interaction x status with measured lift, and the skill tag is MANDATORY and
enforced -- untagged entries are refused, and PROVEN without sufficient
evidence is refused. validatedSkills() returns the coherent profile as it
stands, which is {} for every archetype, by design.
BOMBER x hits conditioning was tested across the order's categories and every
result is underpowered: arsenal (barrel x breaking share) incremental +0.043,
batted-ball (launch x pitcher GB) +0.001, contact quality -0.020 and -0.015,
K x K -0.063. Within BOMBER the counter still leads on hits, 0.218 to 0.160,
consistent with the closed pooled negative.
One bug fixed mid-run: fromStatcastRow maps percentage and raw fields only and
does not carry pitch_mix, so the arsenal category first reported n=0 for every
row -- it was measuring nothing rather than failing. Without catching it,
"arsenal doesn't matter" would have been recorded from a column that was never
populated.
On defense: I looked for a derivable proxy before calling it unsourceable, and
there isn't one. We ingest no fielding data at all, and opposing pitchers'
hits-allowed conflates pitching with defense, so it would validate the wrong
skill. It needs Savant's fielding endpoint -- free, same host as the five
feeds already ingested -- and it is not sourced here, because sourcing it to
test at n=282 would answer nothing.
Nothing proved, so nothing was recalibrated and nothing shipped.
4,221 tests green (335 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Two premise corrections first. Pitcher stuff features have NOT proven solo
through the gate -- every one was refused on sample (n=57 against 500). Four
exceed the effect-size bar (arm angle -0.250, whiff +0.213, k rate +0.206,
chase +0.195), which is why they are worth pursuing, but clearing one of three
thresholds is not passing. And the carrier was not blocked only on the lineup
input: that input was built and measured last session at 94.7% coverage. What
blocks it is n, and n was being throttled by the grading cap.
RUNG 1 IS DERIVED AND COSTS NOTHING. Opposing-team K-rate comes from joining
the opposing roster to the batter k_pct values already in statcast_aggregates
-- no new feed. The improvement this session is that it is PA-WEIGHTED: an
unweighted roster mean counts a 12-PA callup the same as an everyday starter,
which is not the lineup a pitcher faces.
That change alone reversed the term's sign. Unweighted, the lineup term HURT
the model (0.1738 -> 0.1285). PA-weighted, it HELPS (0.1738 -> 0.1953). Same
hypothesis, same data -- the derivation was the problem, not the signal, which
is the entire argument for deriving the best honest version before sourcing
anything. Head-to-head is now +0.2592 with a CI of [-0.0167, +0.5645], very
nearly excluding zero, at n=57.
Within archetype, the two strata come out with OPPOSITE signs -- FLAME
incremental -0.152, non-FLAME +0.145 -- and the pooled value (+0.077) sits
between them, which is the shape a conditional effect makes and is invisible
when pooled. That is what stratifying was for. But n is 20 and 24, the
standard error on a correlation there is about 0.22, and the direction
contradicts the theory that predicted a stronger effect for finesse arms. It
is recorded as a structure to re-test, not as a finding.
Rungs 2 and 3 are NOT triggered. A rung fails only once it has been fairly
tested, and Rung 1 is n-blocked rather than failed. Sourcing confirmed lineups
now would be paying for precision on top of a proxy we have not yet measured.
THE RESULT THAT DECIDES THE TIMELINE: yesterday's cap raise is fingerprinted
in production at 907 grades per snapshot, up from 334, with strikeouts going 6
to 17. That puts n>=500 for pitcher Ks about a week out instead of three
months. Operational note: the manual internal snapshot endpoint now 524s at
the Cloudflare edge because grading the full board exceeds 100s -- the run
still completes server-side (this very snapshot was written by a 524'd
request) and the cron is in-process, so a 524 there is not a failure.
Nothing proven, nothing calibrated, nothing shipped. The counter remains
anti-predictive on strikeouts at -0.064 and the skill model leads it by 0.26.
4,221 tests green (335 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Strikeouts are NOT proven -- n=57 against a bar of 500. But the finding that
matters is not a correlation.
THE CAP. Measured on the live slate via the refusal diagnostic: 1,244 unique
gradeable props exist, the 500 cap graded about 334, and because dedupeProps
takes first-row-wins in FEED ORDER, what survives is decided by feed position
rather than value. Pitchers are 2.6% of a batter-dominated feed, so we were
grading SIX strikeout props a slate against 32 available -- putting n>=500
three months away for every pitcher stat. Pitcher props were never being
refused (graded 5, refused 0, suppressed 0); it was truncation.
Raised 500 -> 1500 on measured cost: 721ms per prop at concurrency 5 is about
179 seconds for the full board, against a cron that runs five times a day and
a fire-and-forget caller that never holds an HTTP response. statsapi is free
and unlimited. Concurrency stays at 5 -- one variable at a time. This unblocks
every n-blocked stat in the programme, not just pitchers.
THE ENGINE. pitcherEngine.js is its own engine, not the batter engine pointed
at pitchers: the batter model asks whether contact becomes a hit and reads
contact quality, the pitcher model asks whether the plate appearance ends
without contact at all and reads stuff. Archetypes are FLAME (whiff-led),
SCALPEL (chase-led), SINKER (pitches to contact) and DEFAULT, and a test
asserts the weight keys are not the batter engine's. The projection is K% by
log5 against THIS lineup, times batters faced, through a binomial. An
unclassifiable arm gets the balanced map, never a guessed archetype.
THE MEASUREMENT, at n=57 and contaminated. Four solo features clear the 0.15
effect bar and fail only on sample: arm angle at -0.250 -- the largest
correlation measured anywhere in this programme -- then whiff +0.213, k rate
+0.206, chase +0.195. The batter cluster's best was 0.135. Head to head,
pitch-v1 resolves 0.1285 against the counter's -0.0639, delta +0.192 with a CI
spanning zero.
That negative is the interesting number. The counter is ANTI-PREDICTIVE on
strikeouts: counting a pitcher's recent Ks is worse than useless, because his
recent totals track which lineups he drew and how long he was left in rather
than his skill. It is the one stat where the incumbent has no defensible edge.
A bug caught on the way. resolveTeam wants an abbreviation and the game log
supplies full team names, so the roster join silently resolved nothing and the
first run reported 0% lineup coverage -- the theorized stuff x lineup carrier
was never being tested, not failing. Fixed; coverage is now 94.7%. The carrier
still shows no incremental signal over whiff alone, and adding the lineup term
lowered head-to-head resolution, which is recorded rather than dropped.
Calibration was not reached: nothing passed the first bar. The batter model
and the counter are byte-identical, verified by diff.
4,221 tests green (335 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PREMISE CORRECTION FIRST, because it defines the bar. total_bases has not
passed BAR 1. Its head-to-head is inconclusive at parity -- delta +0.004 to
+0.007 with a CI spanning zero -- and it is contaminated, and no feature of
its passed the gate. It was described last session as the first challenger
that did not LOSE, which is not the same as proven. If it is installed as the
frozen proven reference and every other stat is held to "the identical bar
total_bases cleared", the bar becomes "be inconclusive at parity" and the
whole cluster passes on a null result. The proven set is EMPTY.
HITS IS NOW A FINAL ANSWER. At n=803 it clears the gate's sample requirement,
so its features were properly TESTED rather than refused: every one fails on
effect size (max marginal |r| 0.053 against a 0.15 bar), every interaction's
incremental contribution collapses to about zero, and the model loses
head-to-head by 0.096 with a CI excluding zero. That is a well-powered
negative and hits should be closed rather than retried.
The rest are n-blocked: total_bases 383, rbi 391, home_runs 228, runs 188,
against a bar of 500. Two leads are worth carrying. home_runs barrel rate has
a marginal r of -0.135, and the sign matters -- higher barrel rate goes with
the counter OVER-predicting, which would be a correction rather than a new
predictor. And runs batterK x pitcherK has the largest incremental in the
cluster at +0.132, with a clean mechanism: strikeouts destroy plate
appearances, and a PA that never happens cannot score.
RBI deserves a caveat rather than a verdict. It is power times OPPORTUNITY,
and we ingest no baserunner state at all, so half its mechanism is missing. A
weak RBI result is evidence that we are modelling half the stat.
total_bases was held frozen: git diff on skillProjection against the prior
commit is empty. The counter is untouched.
Also fixed and verified in production: the point-in-time retention shipped
after yesterday's refresh had already run, so statcast_history was empty, and
its first run then failed on a hand-enumerated schema that had already drifted
from its source ("could not find the 'swing_pct' column"). The refresh itself
still succeeded and wrote all 1,387 aggregate rows, which confirmed the
best-effort guard in prod. The table now mirrors the source via LIKE and the
writer passes rows through whole. Verified live: 1,387 rows retained at as_of
2026-08-03. A usable point-in-time window starts 2026-08-04.
Stage B has nothing to calibrate. Everything now waits on a point-in-time
window and on sample -- both waiting problems, not building problems.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The first production run of the point-in-time retention failed with
"Could not find the 'swing_pct' column of 'statcast_history'" -- the
hand-enumerated column list had already drifted from the table it was copying.
The refresh itself still succeeded and wrote all 1,387 aggregate rows, which
verified the best-effort guard in prod: a retention failure does not fail the
refresh.
The table is now created with LIKE statcast_aggregates, and the writer passes
the row through whole instead of hand-stripping columns, so there is no drift
surface left. Recreating was safe -- nothing had been retained.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Nothing passed. Nothing promoted. Counter byte-identical.
THE BLOCKER, which is the real finding. statcast_aggregates is upserted in
place and holds exactly one as-of date. Yesterday's skill backtest was honest
only by accident: the nightly refresh was unreachable code, so the profiles
sat frozen at 2026-07-21 -- before the settled window. Repairing that cron was
right for production and it refreshed them to today, destroying every prior
version. Scoring a 2026-07-25 game now uses a season aggregate that contains
that game. Point-in-time validation is structurally impossible from that
table, so every number in this run is contaminated and directional, and none
of it is a gate verdict.
Fixed forward: statcast_history retains a dated snapshot on every refresh, so
point-in-time becomes "as_of_date < game_date, most recent". Retention is
best-effort and cannot fail the refresh; both properties are unit-tested. It
has one day of data, which is not yet a window.
SOLO BASELINE, n=383, Bonferroni across 12 tests (alpha 0.00417): nothing
passes. hard_hit_pct is closest at marginal r 0.135 with p 0.0080, failing
both the 0.15 effect bar and the corrected alpha. And it drifted DOWN from
0.153 at n=295 -- an estimate regressing as noise averages out, not an effect
firming up. I called that number encouraging yesterday; on 88 more rows it is
fading, and it should not keep being quoted at its best value.
INTERACTIONS, each scored by partial correlation against the counter residual
controlling for both of its own components: none pass. Only barrel x power
archetype has an incremental exceeding its parts (-0.101 against 0.019) at
n=260 -- the shape Discipline 2 predicts, but a lead, not a finding.
A methodological catch worth keeping. The archetype conditioner was first
built as barrel_pct over league barrel -- a monotone transform of one of its
own components -- so the "interaction" was barrel squared, measuring
nonlinearity in barrel rate rather than any archetype effect, and it produced
this run's only positive result. A Gauss-Jordan pivot test does not catch that,
because the two columns differ by a scale factor. Fixed with a scale-free
collinearity check plus real archetype labels joined from model_snapshots.
Without it this document would have reported a fabricated interaction as the
session's finding.
COMBINED vs COUNTER on total bases: 0.2718 against 0.2647, delta +0.0071, CI
[-0.065, +0.079] -- inconclusive, and the first time a challenger has not
lost. The same engine on hits was -0.116 with a CI excluding zero. That
contrast is the whole argument for total bases, and it is what the physics
said: contact quality governs extra bases, not whether a grounder finds a hole.
Also built: the compound TB projection. skillProjection no longer refuses
total bases -- a deterministic bases-per-hit multiplier had made P(TB>=2)
exactly P(hits>=1), a relabelled hits curve. It is now a convolution over
per-PA base outcomes with hit-type shares shifted by skill. Non-degeneracy is
locked by test.
4,204 tests green (334 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PREMISE CORRECTION FIRST. statModel.js and correlateValidator.js do not exist
in this repository. The validation spec's only prior form is
src/services/python/blueprints/unconventional.py -- a Flask blueprint in the
Python service that is offline in production, scoring NBA factors against a
warehouse that was never populated -- and tests/unit/supplementSystems.test.js
requires only fs and path while defining its own validateFactor inline at line
368. Those tests assert a re-implementation of the thresholds, not an
implementation, which is exactly why they passed for months while nothing was
connected. The diagnosis behind the order is right -- every challenger was
measured without a gate -- but the cause is that there was no gate on the Node
side to import. So it is built, to the exact spec.
correlateValidator: n>=500, |r|>=0.15, p<0.05, Bonferroni across the sweep.
The p-value is exact rather than approximated (t-transform through a
regularized incomplete beta) and is verified in the suite against known
values, because scipy is not available here. Pairs with an unknown side are
dropped, never zero-filled -- a zero-fill inside a correlation does not add
noise, it invents a point at the origin.
THE RUN, hits, n=570, Bonferroni-8: every skill feature fails, and not
narrowly. The strongest marginal correlation against the counter's residual is
0.062 against a 0.15 bar. That is an effect-size failure at a sample that
would have found a real effect comfortably -- a clean, well-powered negative.
The head-to-head agrees: value engine 0.0499 against the counter's 0.166,
delta -0.116 with CI [-0.189, -0.043]. Not promoted.
THE RUN, total bases, n=295: cannot be tested, and that is the finding.
hard_hit_pct shows a marginal r of 0.153 -- above the threshold -- and exit
velo 0.124, refused solely because n is 205 short of 500. It is the most
encouraging number this work has produced, and it is what the physics
predicts: contact quality governs extra bases, not whether a grounder finds a
hole. We have been testing skill inputs on the one stat where they should not
matter much.
Two things the run forced. Feature verdicts are now PER STAT, because marking
these DEAD sport-wide on hits evidence would have killed, for total bases, the
features that look most alive there -- per-sport doctrine one level deeper.
And the gate now reports r and p even when underpowered, because "not enough
data yet" and "nothing here" demand opposite decisions and a bare refusal was
hiding the best signal on the board.
Next: build the compound TB projection (skillProjection still refuses total
bases by design, since a deterministic bases-per-hit made P(TB>=2) identical
to P(hits>=1)), accrue to n>=500, re-run this gate. Leave hits alone.
4,200 tests green (334 suites); web build exit 0; counter byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Built src/services/model/ -- the forward, archetype-selected, skill-based
projection, as a challenger. The champion is untouched.
featureRegistry makes "earn its place or it's out" structural rather than
aspirational: CANDIDATE / PROVEN / DEAD per feature per sport, liveFeatures()
returns PROVEN only, promotion requires n>=200 with positive lift and a CI
excluding zero, and there is deliberately no override argument. It ships with
exactly ONE proven feature -- the incumbent counter, because it is the only
one with a measurement. A test asserts that with only PROVEN features allowed
the projection returns null, so an unproven model cannot reach a user by
accident. The three champion adjustment layers are registered DEAD with their
reasons so they cannot be silently rebuilt.
skillProjection is a PA outcome tree: K and BB combined by log5 odds-ratio
against league (both identities unit-tested), then archetype-weighted contact
quality against contact allowed, then Binomial(PA, p_hit) mixed over a PA
distribution. Archetype is a FEATURE SELECTOR, not a nudge -- BOMBER reads
barrels at 0.50 and ground-ball speed at 0.00, GHOST inverts it -- and a test
locks that the same hitter read two ways moves more than 0.15.
STAGE A: IT LOSES. Out-of-sample on 570 settled hits props with 91.9%
opposing-pitcher coverage, resolution 0.0499 against the champion's 0.166,
delta -0.116 with CI [-0.189, -0.043]. It is not selective either: its eight
most confident picks hit 50%, a lift of -0.065. Not promoted. The gate did its
job on its first real test, which is the point of having built it that way.
Two false starts, both recorded because they nearly produced a wrong verdict:
statcast_aggregates stores PERCENTAGES, so raw rows made bip = 1-29.6-17.1 and
refused 568 of 576 -- the honest-absent guards made a units bug loud instead of
silent, and the conversion now lives at one chokepoint. And the first run
resolved an opposing pitcher for 1 of 570 rows, because ledger team/opponent
are NULL, so it would have reported "skill-v1 loses" while measuring a
batter-only model with no matchup in it at all. The verdict above is from the
corrected run.
The loss is real but partial: park was passed as 1.0, handedness and
opportunity_drift never fired, PA is season-PA over a constant, and the skill
profiles carry no recency at all while the champion has a last-5 term.
Also fixed: the Statcast nightly refresh was unreachable code. It sat inside
tick() below "if (!HOURS_UTC.includes(h)) return" while testing h === 11, so
it had never run once; the aggregates were 13 days stale and both of its
alerts were in the same dead branch. It now runs on its own tick, and the test
that passed happily throughout -- it only checked the string existed -- is
replaced by one that asserts it is not behind the guard.
4,182 tests green (333 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
READ-ONLY. src/ and web/ untouched.
Inventoried every forward-model component against the real objective -- a
forward matchup projection, not an edge number. The finding is that all of it
already exists and is already loaded in production, and 100% of it sits
DOWNSTREAM of the grade in challenger columns nothing serves. The served p_win
reads three features and a game log; it has never seen a pitcher.
Inputs are HAVE, not missing: statcast_aggregates carries 1,354 rows (750
pitchers, 604 batters) with exit velo, launch angle, barrel, hard-hit, whiff,
chase, pitch mix, GB/FB, arm angle, and handedness complete on every row. Real
gaps are team defense and catcher/umpire. So Stage A is a plumbing-and-
modelling job, not a data-acquisition job.
Found along the way: the Statcast nightly refresh is unreachable code. tick()
returns for any hour not in HOURS_UTC (14,19,22,1,3) and the refresh block
then tests h === 11, which that guard can never admit. The mechanism data has
been frozen at its 2026-07-21 backfill for 13 days, and the block's own
failure alert sits in the same dead branch -- the identical silently-guarded-
out shape as the settlement outage.
Design shows the counter: every factor label the SIGNAL BREAKDOWN renders is a
restatement of recent frequency (l5_hot_vs_line, l20_over_line, back_to_back,
home_game) plus several structurally-NBA labels (referees, coach pace,
starters out) inside a baseball product. Not one names a pitcher, pitch type,
handedness or park. The card's forward-read slots already exist and go
unfilled -- the surface needs feeding, not redesigning.
On what changes: the prior measurements were outcome-accuracy, not edge, so
the metric was right and the question was narrow. proj-v1.1 and hits-v1 stay
correctly refuted as DISTRIBUTION swaps on thin inputs -- neither tested a
matchup-fed projection. arch-v1 is a market-relative nudge by construction and
is the one component genuinely measured on the wrong axis. AT CEILING is
provisional: measured only against features the champion already reads.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
READ-ONLY. src/ and web/ untouched; 4,159 tests still green.
WHAT THE CHAMPION IS. probabilityEstimator is five lines of arithmetic: the
empirical frequency of (stat > THIS line) over the game log, blended 0.6/0.4
with the last-5 frequency, then +/-0.03 opponent, +/-0.015 home/away, a
cv>0.40 pull toward 0.50, and a clamp to [0.10, 0.95]. It reads three
features. featureCache retains a dozen more that p_win never touches.
THE ABLATION IS EXACT, NOT A REFIT. Every adjustment is closed-form from
stored features and the consistency step is linear, so each layer subtracts
algebraically out of the stored p_win -- no re-estimation, no re-fetch, no
lookahead possible. Per stat, paired bootstrap:
removing ALL THREE adjustments changes resolution by NOTHING on every stat
hits -0.0059 total_bases -0.0015 rbi +0.0106 runs +0.0130 walks +0.0008
and rbi's home/away is mildly HARMFUL (+0.0053, CI excludes zero). So ~100% of
the champion's resolution is base+recency: how often this player has cleared
this number lately. Everything else is decoration.
A CORRECTION. Pooled, the champion resolves 0.46; per stat it is 0.196 (hits)
to 0.499 (rbi). Pooling stats with different base rates inflates correlation,
so 0.46 should not be quoted as the champion's resolution. Last session's
paired differences remain valid; only the absolute level was inflated.
THE BIGGEST LOSS IS NOT A MISSING FEATURE -- IT IS THE CLAMP. 358 of 1,741
settled rows (20.6%) sit on the boundary, so the model emits a constant there
and cannot rank a fifth of the book at all. And that constant hides two
opposite failures: 0.900 covers home_runs-under truly winning 99.5% (9.5pts
under-confident) next to hits-under truly winning 51.9% (38.1pts over-
confident). PROB_CEIL=0.95 makes the 99.5% case inexpressible. Global
over-prediction is +3.5pts, +7.6 on total_bases. None of this needs new data.
ONE REAL MISSING-WEIGHTING LEAD: opportunity_drift, residual corr +0.156 on
hits and +0.145 on total_bases -- it REPEATS across independent stats, unlike
the weather hits on TB which sit inside the expected false-positive count (70
tests at alpha .05 expects 3-4). And we already compute it: arch-v1's
opportunity axis uses it and extracts nothing (delta +0.0001). Wrong
implementation, not a missing feature -- opportunity must scale the rate, not
nudge the probability.
ARCHETYPE IS UNMEASURABLE, NOT REFUTED. Only 2 of 41 labels (BOMBER, GHOST)
reach n>=40 settled rows and every mean residual straddles zero. That is "we
have not measured it", and it does not license acting in either direction.
Why every challenger has failed is now legible: the ladder and hits-v1 REPLACE
the frequency question with a fitted distribution; the environment axis adds
inputs the champion ignores. Asking the frequency question at the traded line
is the thing that works.
Flagged, not fixed: model_snapshots.outcome is NULL on all 22,032 rows -- the
retention table built for exactly this replay was never settled, so labels had
to be joined from ledger_entries.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The prod-write fingerprint that was blocked by the odds outage has landed on the
first snapshot after deploy. hits-v1 records exactly as the live-board
verification predicted, and the takeable axis behaves as specified -- scope is
book identity, never price shape.
The verdict is unchanged: hits-v1 is REFUTED and stays unpromoted. This confirms
only that it is recording, so the forward accrual can judge the backtest.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
PROMOTE-THE-EARNED. Nothing was promoted, because nothing earned it -- not
because the bar was held high. Measured on the same bar that refuted hits-v1:
own rows only, direction-aligned, paired bootstrap, promote only on a CI
excluding zero.
arch-v1 n=1741 delta 0.0000 CI[-0.0050,+0.0054] inconclusive
contact-v1 n=1055 delta +0.0008 CI[-0.0052,+0.0069] inconclusive
proj-v1.1 n=1664 delta -0.0301 CI[-0.0543,-0.0060] reliably WORSE
matchup/tb-v1/hits-v1 n=0 genuinely pending (rows dated 08-02+)
arch-v1 is the interesting one: it MOVED 76% of rows by 2.5 points on average
and resolution is identical to the champion to four decimals, on the moved
rows too. That is active movement carrying no information -- a finding, not a
pending verdict.
These are true prospective holdouts: arch-v1 and contact-v1 wrote p_win at
grade time into their own columns before the game. Nothing recomputed.
THE 429, read-only. The premise was that we re-pull the full picture every
slot and blow the quota. Measured: PropLine is at 5 calls of 3,000/day --
0.17%. One snapshot is ONE PropLine call per sport, all markets comma-joined.
There is no request-pattern problem, so a change-based pull cannot fix it and
no tier upgrade is needed.
The 429 is odds-api: 478/500 MONTHLY, blocked at 95%. oddsService falls
through silently when PropLine returns empty, and the backup's quota gate
throws the error -- so an empty slate is indistinguishable from an outage and
the message names the wrong provider. Flagged for its own order.
Could NOT verify PropLine movement endpoints: docs are auth-gated and the keys
are production-only. Not asserted either way. The movement-as-data argument
stands on its own merits and should be justified that way, not as a quota fix
it isn't.
Book-breadth invariant written down: we never discard books. All are kept and
shown (DISPLAY_BOOKS = MODEL + REFERENCE + DFS); DFS pick'em is excluded from
PRICING only, because a fixed-payout shaded number is not a market price.
Verified this is already what bookRoles.js does.
Champion byte-identical; every challenger stays wired.
4,159 tests green (332 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The self-learning loop stopped two days ago and reported success the whole
time. 1,444 ledger rows from 2026-08-01 sit unsettled with settle_attempts=0
-- never even attempted -- and every accruing challenger has been starved of
settled sample as a result.
ROOT CAUSE. settleLedger fetched open ids, then REFETCHED the full rows with
.in('id', ids). PostgREST puts filters in the URL, so 500 UUIDs became an
18,499-character request that the fetch layer rejects with "TypeError: fetch
failed". The result was destructured as `const { data: rows } = ...` with NO
error binding, so rows came back null, the loop body never executed, and the
function returned {settled:0, voided:0, unrecoverable:0, pending:0} --
byte-identical to a clean "nothing to settle". Reproduced against prod before
changing anything.
WHY IT HID FOR TWO DAYS. It is volume-triggered. Daily volume ran 20-260 rows
and settled perfectly for weeks; 2026-08-01 was the first day past the 500-row
fetch limit. And the zero-settle ops alarm reads these very return values, so
pending:0 told the watchdog the backlog was empty -- the alarm built to catch
exactly this could not see it.
THE FIX. The refetch existed only to add game_date/settle_attempts/
dclv_computed_at. Selecting them in the first query removes the id list
entirely, so there is no URL to overflow at any volume. A failed fetch now
surfaces its error instead of being reported as an empty backlog.
captureClosing carried the same shape one level down -- .in('id', g.ids) on an
UPDATE, which fails identically once a single line|odds group gets large on a
big slate. Its id filters are now chunked at 100 (~3.7 KB).
Tests: the regression is locked by asserting settlement issues NO id-list
filter at 500 rows, and that a failed fetch is never reported as an empty
backlog -- the two properties that would have caught this. Two existing
suites asserted the old two-query shape and were updated to the real one.
4,159 tests green (332 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
The prod-write fingerprint did not land: the odds provider is returning 429
(quota exhausted), so the snapshot refuses with gradeCount 0 and the MLB board
has been frozen since 07:30 UTC. The 14/19/22 UTC cron slots failed the same
way, all before this change deployed -- hits-v1 sits inside the snapshot's
existing try/catch, is purely additive, and had zero grades to attach to.
Firing is already verified against the real production snapshot through the
real attachProjection path (158/159). What is pending is only confirmation
that the deployed process writes the columns, which needs a slate the pipeline
can fetch. The exact fingerprint query is recorded in the spec.
The odds quota exhaustion is a live outage of the whole grading pipeline and
is flagged for its own order, not folded into this one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
Hits was diagnosed as a family mismatch: 84% of hits rows trade at 0.5, so
the stat rides on P(0), and a negative binomial has unbounded support and no
notion of opportunity at all. hits-v1 models it as the bounded conversion it
is -- N ~ the player's empirical at-bat distribution, hits|N ~ Binomial(N,q),
with the multiplier scaling q (conversion) and never N (opportunity).
STEP 0 confirmed the inputs before the model existed: 30/30 real ledger
players, 100% combined-input coverage. Every read goes through knownRate --
a row with no atBats is dropped, never counted as a 0-at-bat game.
It FIRES: 158/159 hits props (99.4%) on the live production snapshot, through
the real attachProjection path. Scoping by book IDENTITY rather than price
shape kept 94 out-of-promotion-band props on the board, 93 of them modelled --
59% that a price rule would have deleted.
And it LOST. Point-in-time replay (game log truncated strictly before each
row's game_date, real grade-time multiplier), hits-only, direction-aligned,
n=242: resolution champion 0.195 / ladder 0.048 / hits-v1 0.026. Paired
bootstrap on the same rows: hits-v1 - ladder = -0.022, CI95 excluding zero.
Not promoted.
The value is in what it eliminates. The family was wrong AND the mean was not
the constraint -- hits-v1 moved the line-0.5 mean 0.554 -> 0.581 toward a
0.598 base rate while resolution fell. What is left is per-prop
discrimination: the ladder's inputs, not its distribution.
The pre-registered fallback is recorded as WRONG rather than deleted. It said
hits might be genuinely low-resolution for anyone; the champion scores 0.276
on the identical 189 rows, so there is real signal and the ceiling claim was
the comfortable reading, not the honest one. Its own control refuted it, and
that control was already in hand when the branch was written.
hits-v1 stays wired as a challenger writing its own ledger columns so the
forward accrual can confirm the backtest. Champion, ladder, ranking,
calibration, reference ruler and the four accruing verdicts are byte-identical
-- the diff has zero deleted lines.
Tests 4,156 green (332 suites); web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
BYTE-IDENTICAL. The audit found no consumer getting the wrong axis, so this
is a disambiguation, not a bug fix. 4,131 tests / 331 suites green.
STEP 1 AUDIT -- and the order's premise was wrong in a useful way:
the four accruing challengers read the flag ZERO times (not four)
the ranking gate wants PROMOTION, gets promotion [correct]
the ledger column holds the LEDGER band, consumed as such
the UI (LiveHeroProp) TYPES a `takeable` field it never renders
THERE ARE THREE DEFINITIONS, NOT TWO -- and I only found the third by
tracing the ranking gate:
1. IDENTITY can it be bet? book identity (takeability)
2. LEDGER BAND worth recording? odds >= -160, UNCAPPED plus
3. PROMOTION worth crowning? -160..+200, i.e. band PLUS a ceiling
(2) and (3) genuinely disagree, and I measured it rather than asserting it:
439 rows -- 28.2% of all takeable=true ledger rows -- carry prices above
+200, up to +1300. A +1300 longshot is a real bet worth RECORDING and not
one worth CROWNING. Both are correct for their own purpose.
THE DANGER WAS NEVER THE LOGIC. It was that three questions shared one
word, so a reader could not tell which answer they held -- and hits, which
must model thin/juiced/one-sided REAL markets, would have been the next
reader to guess wrong.
RESOLUTION: all three now have distinct names in config/takeability.js;
gradeRanking calls isWithinPromotionBand so its intent is self-evident (a
test pins it byte-identical to the old valueEngine call across the whole
price range); the ledger dual-writes within_price_band with `takeable`
kept as a documented DEPRECATED MIRROR so nothing breaks. Column comments
in the database now say what each column actually holds.
I did NOT redefine `takeable` in place. Four readers and a ranking gate
sit on it, and silently changing its meaning under cover of a naming
change is exactly the class of move this session keeps removing.
Gates: 4,131 tests / 331 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Both guards are ADDITIVE. The full suite (4,111 -> 4,126 tests, 331 suites)
passes unchanged through the migration, which is the evidence that no
currently-correct output moved: served path, champion, reference ruler and
the four accruing challengers are byte-identical.
GUARD 1 -- src/utils/known.js. Number(null)===0 has produced at least SIX
separate defects here, including one in a module written the same week its
author documented the trap. Per-module vigilance has demonstrably failed,
so the rule lives in one place and SEVEN sites now delegate: platoonSplits,
projectionChallenger, challengerProjection, contactChallenger,
statcastAggregateService, consensusRuler, gradeRanking -- plus
compoundTotalBases moved onto knownRate.
Two functions, deliberately: knownNumber (any finite number -- a REAL 0 is
a fact and must survive) and knownRate (non-negative, rejects booleans --
for counts/rates where `true` or -1 is broken, not thin). Collapsing them
is how the next variant gets in. firstKnown() exists because `a || b`
discards a measured 0 and `a ?? b` does not.
MY OWN GUARD HAD THE BUG IT EXISTS TO PREVENT, and its own test caught it:
Number([]) === 0, so an empty array coerced to a measured ZERO. Same trap
wearing a different type. Both helpers now reject objects outright.
GUARD 2 -- src/config/takeability.js. Takeability is BOOK IDENTITY and
never price shape. Baseball prop markets are genuinely thin, juiced and
one-sided, and all three are NORMAL structure: betrivers and hardrockbet
legitimately quote one side only (5 such rows surfaced in yesterday's
re-stamp), and a hits-over at -300 is a real placeable bet. A rule that
inferred un-takeability from price extremity or one-sidedness would throw
those away while still admitting a DFS book at an ordinary -119 -- exactly
backwards, because the -119 is the fake one.
THE DISTINCTION THAT MUST NOT COLLAPSE, now enforced by test:
isTakeableMarket(book) -- CAN it be bet? (identity)
isWithinPriceBand(odds) -- SHOULD we promote? (policy band, floor -160)
A -300 DraftKings prop is takeable AND out of band; a PrizePicks -119 is in
band AND not takeable. Independent axes.
FLAGGED, NOT SILENTLY CHANGED: the ledger's `takeable` column is the
PRICE-BAND answer, and its name predates this distinction. Four challengers
and the ranking gate read it, so renaming or redefining it is its own
order -- doing it here would have changed correct current behaviour under
cover of a hardening change.
Fixtures are REAL prod rows from the 2026-08-02 re-stamp, not invented.
Gates: 4,126 tests / 331 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Database only; no application code changed, so the served path, champion
and reference ruler are byte-identical.
RESULT: 862 rows re-stamped from the takeable LOCK-TIME price in
lock_lines, 862/862 now anchored to takeable books, tagged
price_source='archive_restamp', quarantine lifted. 812 pending clean rows
recovered into the accruing verdicts. Holdout verification: 2,792 rows,
862 re-stamped included, 0 re-stamped rows non-takeable, 144 still
excluded, 0 quarantined rows leaked, and 0 NON-TAKEABLE rows remain in the
holdout population since 2026-08-01.
DEVIATION, stated rather than buried: the order authorised 936. That
figure came from a takeable book posting the same LINE. Requiring what a
re-stamp actually needs -- that book's price for the GRADED SIDE at LOCK
TIME -- resolves 862. Of the other 74, 73 have a takeable side-price only
OUTSIDE the lock window and 5 are genuinely one-sided markets.
I did not widen the window to reach 936. A takeable price captured hours
after the grade is a later market moment, not a lock price; substituting it
is precisely the reconstruct-vs-join line this order was fenced against,
and it would have been invisible in the totals -- showing only as a
cleaner-looking 936.
Those 74 were also RE-TAGGED, because their old label had become a lie:
recoverable_same_line -> no_takeable_lock_price_for_side. A future attempt
reading the old tag would have been invited to widen the window and call it
recovery.
takeable was RECOMPUTED from the recovered price rather than carried over
-- the old flag was computed FROM the contaminated price and was wrong on
its own terms. 101 rows had their flag change, which is the direct measure
of how wrong it was.
Provenance travels with the data (price_source), on the same principle as
is_proxy: a value recovered by a later join is not identical in kind to one
captured natively at grade time, even when it is the same number.
EVIDENCE FOR THE NEXT ORDER'S INVARIANT: 5 of the excluded rows are
one-sided TAKEABLE markets, and betrivers/hardrockbet legitimately quote
one side only. A guard that inferred takeability from price shape would
throw away real markets while still admitting a DFS book at -119 --
takeability is book IDENTITY, never price extremity or one-sidedness.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
PART 1 verified by inducing the REAL rowsFromSnapshot over REAL lock_lines
rows from prod. Three cases, 0 non-takeable anchors:
Narvaez (dabble/kalshi/prizepicks/smarkets, NO takeable book)
-> book=null, price=null, takeable=null [honest absent]
Schwarber(bovada/dabble/novig/PINNACLE before draftkings)
-> draftkings +102 [pinnacle SKIPPED, proving TAKEABLE not MODEL]
Ohtani (dabble/onexbet before draftkings) -> draftkings -266
Narvaez is the case that matters: pre-fix he was stamped dabble +104
takeable=true; he is now honestly absent.
A HARNESS BUG RECORDED: my first verification pulled live /api/odds/mlb,
which returned {"error":"Odds data temporarily unavailable"}. The script
read that as 0 props and printed "all from takeable books? true" -- a
VACUOUSLY TRUE pass. I caught it only because I also printed the book list
and it was empty. Same family as the silent-false traps: a probe that finds
nothing looks identical to a probe that finds nothing wrong.
PART 2: 1,006 rows tagged via the purpose-built quarantine_reason at ROW
level with three sub-cases (recoverable_same_line 936, no_takeable_quote
49, takeable_line_differs 21). getModelAggregate ALREADY excluded
quarantined rows, so the public record and the n>=20 gate were clean
automatically; all five committed holdout scripts now carry the exclusion
explicitly.
PART 3 -- the re-stamp call is now fact-based. The takeable LOCK-TIME price
is recoverable for 936/1,006 (93.0%) from lock_lines, the correct
instrument. Only 431 appear in closing_captures, which is the wrong timing
for a lock price anyway.
LINE CONTAMINATION ANSWERED (previously unverified): the stored line
MATCHES a takeable book's line on 936 (93.0%), DIFFERS on 21 (2.1%), and is
unverifiable on 49 (4.9%) where no takeable book quoted the prop at all.
That makes it cleanly row-level: re-stamp the 936 as an honest JOIN and
recover 886 pending rows for the holdouts, or leave all 1,006 excluded.
Either way the 21 + 49 stay out -- re-stamping those would invent a lock
price, or a line, we never captured. Nothing re-stamped; Kev's call.
Gates: 4,111 tests / 330 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Ships before tonight's settle. Served path, champion, ranking and the
reference ruler are untouched.
TWO leaks, not one. The audit found ledgerService.indexProps; tracing the
lock price found that snapshotService.indexOdds has the SAME defect -- it
also indexed the full props list, so gradedAt.odds (the price a grade is
locked at) could itself be a DFS or exchange price. Fixing only the ledger
would have left the contamination flowing in through the lock.
Both now gate on TAKEABLE_BOOKS -- deliberately NOT MODEL_BOOKS. pinnacle
is model-eligible and correctly not takeable, so a MODEL gate would
re-break this the moment pinnacle's feed recovers. A test asserts pinnacle
cannot anchor a price.
TWO INDEXES, TWO ROLES, because the row needs two different things from a
prop and they have different correctness rules:
PRICE / BOOK / TAKEABLE -- takeable books only.
GAME FACTS (game_time, game_date, team/opponent) -- book-INDEPENDENT.
First pitch is first pitch whichever book listed it, so these still
come from any book. Gating them too would drop otherwise-valid rows
for no gain.
Collapsing those roles into one index is precisely the bug.
No takeable quote leaves the key ABSENT and the price null. An honest
missing price beats a price from a book you cannot bet -- and it keeps the
takeable flag from being computed off a DFS number, which is what made it
wrong on its own terms rather than merely mislabelled.
Gates: 4,111 tests / 330 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
READ-ONLY. Nothing enforced or fixed; the five challengers untouched.
VERDICT: gaps exist, and one is LIVE CONTAMINATION of the ledger -- the
exact table every accruing holdout resolves against. book, locked_odds and
the takeable flag ITSELF are being stamped from books you cannot bet: DFS
dabble (707 rows, 24% of all rows), offshore bovada (214), onexbet (42),
exchange kalshi (7, mean |odds| 1120).
0% before 2026-08-01. 47.9% on 08-01. 42.5% on 08-02. It began the day I
widened the books for display.
LEAK LOCATED, not inferred: recordPipelineGrades indexes byKey over the
FULL display-widened props list, then prefers that prop -- book:
(prop && prop.book) || g.book, and locked_odds/takeable both fall back to
oddsForSide(prop). The grade is computed on a MODEL book and the ledger row
is then re-stamped from whatever book indexed first. The takeable flag is
therefore not merely mislabelled: it is computed FROM the contaminated
price, so it is wrong on its own terms.
The served grade path is clean TODAY (428 grades, 100% MODEL books), so
dedupeProps' gate works. But MODEL_BOOKS is NOT a subset of TAKEABLE_BOOKS
-- pinnacle is model-eligible and correctly not takeable -- so the
projection may anchor to a reference line by design. Harmless while
pinnacle returns nothing; live again when it recovers.
BLAST RADIUS bounded but growing: 47 contaminated rows have already
settled (21% of settled rows since 08-01) and ~700 are still pending and
will settle into the holdouts. The damage is mostly ahead of us, which is
what makes this urgent rather than historical.
NOT VERIFIED and not claimed either way: whether the stored `line` is also
contaminated. It traces to the graded prop, but I did not check it
end-to-end; the enforcement order should.
The prediction-vs-reference distinction HOLDS and must not be collapsed:
the prediction target must be takeable, while fair_prob / consensus / edge
stay reference. The bug is not the three-way split -- it is that one write
path ignores it.
Stack sequenced in the plan: (a) takeable enforcement, (b) structural
Number(null)===0 guard (hits will re-trigger it -- its 0.5 lines make P(0)
the whole game), (c) hits. Carry-forward: tb-v1 verdict, the third
pre-registered branch, and the 100s Cloudflare timeout vs a ~115s snapshot.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Firing verified on a real prod snapshot: 10/10 total_bases props carry
proj_tb_p_over. The snapshot HTTP call returned 524 (Cloudflare's 100s
origin timeout vs a ~115s snapshot) but the work completed server-side --
confirmed from the ledger rather than assumed.
Face validity is good and diagnostic: means agree almost exactly with the
ladder (1.813 vs 1.833), so this is a SHAPE-ONLY intervention, which is
what was intended. Component rates are plausible, and Carroll's triples
rate (0.112, far above his peers) is a clean check -- he is a speed player
and the model sees it.
AN OBSERVATION I AM NOT RESOLVING BY EYE: tb-v1 reads systematically LOWER
than the ladder (0.424 vs 0.540 at the same mean). That is the expected
DIRECTION, since the NB overstates P(>=2) by treating a home run as four
accumulating events -- but whether 0.424 is right or an overcorrection is
not knowable from face validity. A ~1.8-TB hitter clearing 1.5 empirically
sits nearer 45-50%, between the two. I am not claiming tb-v1 is better; the
holdout decides.
BRANCH PRE-REGISTERED, before the result, so the verdict cannot be
reinterpreted afterward: improves -> family-mismatch HOLDS, similarity
stays off the critical path, hits is next; does not improve -> hypothesis
WRONG and the mean-weakness/similarity branch REOPENS.
Also recorded: I hit Number(null)===0 in my own new module -- a null
component rate treated as a measured zero, the difference between "never
triples" and "we don't know his triple rate". A test caught it. Sixth
appearance of this trap in this codebase, and it caught the person writing
the warnings about it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Current ladder (proj_p_over_line) and champion p_win are BYTE-IDENTICAL.
tb-v1 writes alongside them, on total_bases props only.
STEP 0 -- components confirmed on real data, not assumed. statsapi has no
singles field, but hits - doubles - triples - homeRuns reproduces stored
totalBases EXACTLY on a real 10-game log. So the decomposition is exact,
not an approximation.
THE MODEL. Each component gets its own per-game Poisson rate; TB is their
weighted sum, and the PMF is built by exact convolution rather than
simulated (TB support is small). It inherits the SAME combined multiplier
proj-v1.1 computes, so the two models differ only in STRUCTURE.
Why this is the fix: with identical mean TB of 1.0, a pure-HR hitter and a
pure-singles hitter get P(TB>=4) of 0.221 vs 0.019 -- a 12x difference an NB
on TB alone cannot express, because it treats one home run as four events.
A test asserts that separation, and asserts P(TB>=4) for a pure-HR hitter
equals P(at least one HR) exactly.
INDEPENDENCE IS AN APPROXIMATION AND IS LABELLED AS ONE: a plate appearance
that becomes a double cannot also become a single, so the components are
weakly negatively correlated and independent Poissons slightly overstate
the tail. Closer to the truth than what it replaces; not a solved problem.
HONEST-ABSENT throughout: fewer than 3 usable games, or no derivable
component, returns null and the prop keeps the current ladder value. An
inconsistent row (hits < extra-base hits) is SKIPPED rather than clamped to
zero -- clamping would invent a plausible line out of a broken one.
I HIT THE Number(null)===0 TRAP IN MY OWN CODE and a test caught it: a null
rate passed a naive finite check and was treated as a measured zero, which
is the difference between "this player never triples" and "we do not know
his triple rate". Both tbPmf and tbMean now reject null/''/boolean strictly.
Holdout committed: TB ROWS ONLY (49 of 437 settled -- averaging into other
stats would hide the effect) and DIRECTION-ALIGNED, since the unaligned
comparison is the artifact that accounted for 41% of the ladder's apparent
loss. If tb-v1 does NOT improve, the family-mismatch hypothesis is wrong
and the mean/similarity branch reopens -- recorded in the query header.
Migration applied: proj_tb_p_over + proj_tb_meta, NULL-meaningful.
Gates: 4,104 tests / 329 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
READ-ONLY. Nothing built or fixed; the four challengers untouched.
41% OF THE REPORTED GAP WAS A MEASUREMENT ARTIFACT. p_win is P(graded
side); proj_p_over_line is P(over); 31.4% of settled rows are UNDER-graded,
so comparing them raw measures the ladder backwards on a third of the
sample. Matched + direction-aligned (n=437): 0.252 vs champion 0.352, not
0.108 vs 0.331. The PRODUCT is not making this mistake -- I checked;
projectionChallenger normalises both to the over basis deliberately. The
error was in the measurement.
THE LOSS IS CONCENTRATED. hits (n=245, res 0.060) and total_bases (n=49,
res 0.009) are 67% of rows and carry essentially no signal. Everything else
is fine or better: walks 0.519 vs champion 0.544, runs mean 0.345 vs 0.392,
and on DOUBLES the ladder's mean BEATS the champion's (0.207 vs -0.062).
IT IS THE MEAN, NOT THE SHAPE. On the two failing families the mean itself
carries no signal (0.052, -0.019) against the champion's 0.158 and 0.085.
Where the mean is good the probability is good -- shape follows mean.
A HYPOTHESIS I TESTED AND DISPROVED: prediction compression. I expected
P(>=1 hit) to sit in a narrow band and fail to rank. It does not -- spread
ratio 0.94 overall, 0.80 for hits, 0.94 for total_bases. The ladder has
comparable spread; it is spread in a direction uncorrelated with outcomes.
Recorded because it was a plausible story the data refused.
PRIORS AND PLUMBING CLEAN. proj_factors carries form_rate,
combined_multiplier and breakdown on every row; proj_point 100% populated
with sane centres (hits 0.830 vs line 0.578). Not the environment-style
silent-null failure.
NAMED CAUSE (structural, flagged as hypothesis not finding): the count
model mismatches those two stats. total_bases is a WEIGHTED SUM (1B..HR =
1..4), so an NB treats one home run as four events and mis-states variance
-- and TB has the worst result in the table. hits is BOUNDED BY AT-BATS and
mostly traded at 0.5, so almost everything rides on P(0), the region where
the wrong family hurts most. walks/runs/doubles ARE genuine low-rate counts
and are exactly the ones that work.
FIX BRANCH: targeted per-stat fix for hits and total_bases. THIS REMOVES
THE MLB SIMILARITY BUILD FROM THE CRITICAL PATH -- that branch assumed a
GLOBAL mean weakness, and the mean is fine or better on three of six stat
families. Similarity may be worth building later, on evidence, not on this.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
READ-ONLY. Nothing connected, built or wired; the accruing challengers were
not touched. "Dormant" meant three different things and in no case is the
answer "connect it".
DISTRIBUTION LADDER IS NOT DORMANT. projection/distribution.js is consumed
by projectionChallenger (proj-v1.1), live on every snapshot at 94.2%
coverage (276/293) with 437 settled rows since 2026-07-23. It is a FOURTH
accruing challenger, and it is LOSING: resolution 0.108 vs the champion's
0.331. That verdict is no longer thin.
It is also PER-STAT and doctrine-correct -- nine distinct league priors
(hits 0.90, total_bases 1.45, home_runs 0.15, ...) each feeding a
gamma-Poisson posterior into a negative binomial. Correcting the plan:
§10.3's "single additive index across hits/Ks/TB" is engine1's GRADE, not
this ladder, which made a solved problem look open.
SIMILARITY IS WRONG-SPORT. Zero callers, and its weights are NBA
vocabulary: pace 0.15, referee_tendency 0.06, lineup_context 0.12,
score_state_context 0.05, travel_fatigue 0.08. MLB has no pace and no
referees. Connecting it would be the sport-stubbed-in-on-another-sport's-
template breach, and it would fail QUIETLY -- missing factors are skipped,
so the score would silently collapse onto whatever few dimensions happened
to exist. CONSTRUCT, not connect.
BAYESIAN WOULD REGRESS THE MODEL. Zero callers, and DISTRIBUTION_SHAPES
keys on rbis / runs_scored / strikeouts_batter / outs_recorded /
pitcher_strikeouts / walks_allowed / pitches_thrown -- NONE of which are
live stat keys (S41: they are rbi / runs / outs / strikeouts).
getDistributionShape defaults to 'normal' on an unknown key, so wiring it
as-is would model COUNT stats as Gaussian, silently, on most MLB props. It
is also superseded by distribution.js. Do not connect; retire or rewrite.
DEPENDENCY, inverted: a better mean would help the ladder, but the ladder
is already connected and both would-be foundations are unusable -- so this
is not "connect similarity first", it is "the ladder is live and
underperforming, and strengthening its mean requires BUILDING an MLB
similarity layer that does not exist".
Next-order pointer moved to diagnosing proj-v1.1: the only candidate
already carrying settled evidence, and its diagnosis decides whether the
similarity build is worth doing at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Records the verification that matters: firing measured on a real prod
snapshot rather than inferred. environment 248/293 (84.6%) -- also its
FIRST confirmed ledger write, which the previous session could only infer
-- and matchup 243/293 (82.9%) on tier batter_own_split. Both were 0/634.
Collinearity guard passed at n=243: r = -0.003 vs the projection, +0.074 vs
p_win, +0.003 vs line, -0.068 vs environment, -0.150 vs opportunity. The
axis is not re-encoding recent form. The nudge distribution is also the
right SHAPE -- mean +0.0007, 123 positive / 120 negative -- a balanced
two-sided signal; a one-sided distribution would have suggested a sign or
baseline error.
Plan reconciled in place: arch-v1 condition axes marked firing, three
challengers listed with coverage and their own holdout queries, and the
next-order pointer moved to connecting the still-dormant layers
(similarity, Bayesian, distribution ladder) with archetype_x_archetype as
the named alternative.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The axis was already wired and firing on 0/634 prod rows. Three separate
absences kept it silent, and all three are now joined:
1. oppPitcherByTeam 0 -> the self-origin /api/schedule/mlb/pitchers route
returned nothing in prod. Added the statsapi probable-pitcher hydrate as
a fallback, mirroring the one the schedule step already uses. 29/30
team-sides, one free request.
2. handById 0 -> follows from (1); the batched people call now has ids.
3. bats 0/120 -> batter hand rode ONLY on statcast aggregate rows, which do
not cover the slate. The season player list we ALREADY fetch and cache
carries batSide on 1342/1342, so this is a join, not a fetch.
Switch-hitters ('S') are preserved as-is; platoonSplits decides what to
do with them, not the map.
Verified end-to-end against the live API: opp_declared 29,
pitchers_with_hand 29, batters_with_hand 1342, and a real read --
multiplier 0.966, L vs R, 287 observed PA, weight 0.324 -- composing
alongside environment in one challenger.
FALLBACK LADDER, and a deliberate deviation from the order. Shipped tier:
`batter_own_split` (the hitter's OWN vs-L/vs-R line, regressed toward HIS
OWN overall rate), labelled on every adjustment.
`league_generic` is deliberately NOT implemented. platoonSplits already
handles thin evidence by regressing toward the hitter's own rate, which
covers the thin case per-player; its own doc-comment argues a hitter with
no split evidence should get NO adjustment. A league split applied to such
a hitter models the LEAGUE, not the player -- the doctrine breach the order
itself names in the same step. Adding it would have produced more firing
rows and a weaker signal.
`archetype_x_archetype` is scoped, not built: it needs the opposing
starter classified per game, which is real work and a separate order. The
tier vocabulary is in place for it.
Honest-absent on every join: no starter, no pitcher hand, or no batter hand
-> NO matchup adjustment, never a fabricated neutral. A neutral multiplier
produces no adjustment row at all.
Holdout committed (scripts/matchup-axis-holdout.sql), filtered to
matchup-carrying rows, and it keeps MATCHUP'S OWN nudge visible rather than
only the combined challenger -- arch-v1 composes four axes into one
p_win_challenger, so a combined-only view could not tell which axis earned
the movement, or which one is dragging.
Champion p_win, ranking, calibration, the armed invariant and the two
accruing verdicts are untouched.
Gates: 4,093 tests / 328 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
My first attempt did not work in prod -- team stayed 0/323 after deploy.
I resolved the team name AFTER the hint-confirmation check, but the check
itself reads hit.currentTeam.name, which is undefined because
/sports/1/players returns { id, link }. With a FULL-NAME hint (what
snapshotService passes) neither branch of teamRecordMatchesHint could
match: the name branch had no name, and the abbr branch cannot resolve a
full name to an abbr. Confirmation failed, the team was nulled, and my
later backfill ran on an already-null value.
withTeamName() now backfills the name from the cached /teams list BEFORE
any comparison, and is used at all three confirmation sites plus the
return. Verified against the live API on all four cases: no hint, FULL-NAME
hint, abbr hint -> "Philadelphia Phillies"; WRONG hint -> null.
That last case matters most: a wrong hint must still REFUSE. The
confirmation exists so a namesake collision cannot tag a player to a team
he is not on, which would fabricate opponents downstream. Making the match
succeed must not make it succeed wrongly, and a test locks it.
Gates: 4,087 tests / 327 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Reconciled in place, not regenerated. Next-order pointer now MATCHUP AXIS
with its verified sourcing table, and an explicit note that
SOURCE-LINEUPS-first is NOT needed.
Marked DONE with their evidence: p_win ranking + edge retirement,
calibration DECIDED, MLB isotonic DECIDED (provisional label retracted),
grade cap 25->500 (board 7->365+), book widening, S59 invariant armed,
environment axis repaired.
Records the honest shape of Phase 1: it is further along than the phase
table implied, but mostly because the work turned out to be CONNECTION AND
REPAIR rather than construction -- the ladder question dissolved, the cap
was discarding 95.7% of the slate, and two condition axes were wired but
firing on zero rows.
Carried forward without softening: WNBA is NOT BUILT rather than failed,
and the ruler is MARKET-not-SHARP with PENDING-RECOVERY status until
PropLine answers the Pinnacle question -- not to be enshrined as permanent.
Remaining ~19 orders, ~9 unblocked. The two accruing verdicts are time,
not code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
PART 1 -- PREMISE CORRECTION, then the real fix.
The order said the invariant's blocker was removed because "team is now
populated 416/416". It is not: what became 416/416 is home_team/away_team.
`team` (the PLAYER'S roster team) is still 0/416. Arming the guard off
home_team would compare the prop's game to itself -- always a match, a
permanent no-op that LOOKS armed. That would be worse than leaving it
disarmed, because it would read as a working guard.
The guard is also ALREADY fail-safe by construction (`if (knownTeam &&
gameTeams && ...)`), so Part 1's requirement was met in code all along.
What was missing was the data.
ROOT CAUSE: /sports/1/players returns currentTeam as { id, link } with NO
name, so searchPlayer's `hit.currentTeam?.name` was ALWAYS undefined and
every resolve returned team: null. The id is present on 1342/1342 and the
/teams list (already cached 24h) maps id -> name, so resolving it costs no
new request. Verified: Schwarber -> Philadelphia Phillies, Ohtani -> Los
Angeles Dodgers, Judge -> New York Yankees.
Five tests lock the fail-safe: drops only on a positive not-in-game;
abstains on unknown player team; abstains on unknown game participants;
and a row carrying only home_team/away_team does NOT satisfy the guard --
so the tautology can never be reintroduced.
PART 2 -- MATCHUP SOURCING: BUILDABLE. Measured on tonight's real board
against the free feeds, by VALUE not endpoint presence (the environment
trap: wired and null 634/634):
opposing starter 29/30 team-sides (home 14/15, away 15/15)
pitcher hand 1342/1342 (pitchHand.code)
batter hand 1342/1342 (batSide.code; L 416 / R 848 / S 78)
SHARED DEPENDENCY, and it is the finding: /sports/1/players -- a list we
ALREADY fetch and cache -- carries currentTeam.id, batSide AND pitchHand.
One join unlocks the invariant's input and two of the three matchup inputs
at once. The third (probable starter) comes from the schedule hydrate that
already exists.
So matchup is BUILDABLE and is the next order; SOURCE-LINEUPS-first is NOT
needed. Archetype-level reach on the opposing starter is available too
(the SP resolves to a player id, so the existing classifier applies) --
noted, not built.
Champion p_win, ranking, calibration and both accruing verdicts untouched.
Gates: 4,082 tests / 327 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Report for the audit + fix already committed. Records the two things worth
carrying forward:
1. The environment axis has NOT yet been observed writing to the ledger,
and I am not claiming it has. recordPipelineGrades upserts with
ignoreDuplicates and dedupes on (user_id, player_key, stat, line, side,
game_id) -- correctly, so a re-run never overwrites the original lock.
Today's 429 rows predate the fix, so the axis cannot backfill onto them;
first ledger observation is tomorrow's slate. What IS directly verified
is the resolver (105/120) and the join key (416/416) -- the two things
that were actually broken.
2. Matchup is not fixed and is not claimed as fixed. It needs the opposing
starter and BOTH hands, and the audit shows three separate absences:
oppPitcherByTeam 0, handById 0, bats 0/120. Fixing the pitcher feed
without the hands, or the hands without the feed, still produces an axis
that fires on zero rows.
Also noted: the S59 slate JOIN INVARIANT keys off the same null `team`
field, so it is currently inert too.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
STEP 0 AUDIT -- the "already partly live" premise was half true: the CODE
is wired, the axes are NOT firing. Across 634 graded prod rows the
environment and matchup axes fired on ZERO rows, while 13 archetype axes
fired normally (power 80, swing_miss 69, contact 56, launch 51,
line_drive 43, ...) plus opportunity 142. Ledger confirms it from the
other side: env_multiplier, env_park_base, env_weather_mod, wx_forecast
and env_weather_state are ALL null on 634/634.
ROOT CAUSE, located rather than inferred. A drop-off audit against the
live snapshot: with_team_field 0/120, with_bats 0/120, with_playerId
120/120, oppPitcherByTeam 0, handById 0. `team` is a KEY on every stored
grade and NULL on 416/416 -- so an environment resolver keyed off the
player's roster team could never find a venue, while buildContext sat
there with all 30 teams mapped and 14 weather forecasts resolved and
unused. Coors composes to 1.241 the moment it gets a key.
FIX -- and it is the more correct join, not just a workaround. The park
and the weather belong to the GAME, not to the player's roster team, and
the game rides on the prop from the odds feed. gradeBestSide now carries
home_team/away_team onto the graded row (the legacy grade shape dropped
them), and contextFor joins on the game first, keeping the roster team as
a fallback. This no longer depends on a stats-resolve that can
legitimately fail.
MATCHUP/PLATOON IS NOT FIXED HERE and is not claimed as fixed: it needs
the opposing starter and both hands, and the audit shows
oppPitcherByTeam=0, handById=0 and bats=0 on the slate -- three separate
absences. Per "one axis at a time" that is its own order with its own
diagnosis, not a second fix smuggled into this one.
Champion p_win, ranking, calibration and opportunity_drift's accruing
verdict are all untouched.
Gates: 4,077 tests / 326 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The arch-v1 environment and matchup axes fired on ZERO prod rows across
634 graded props while archetype axes fired normally, and buildContext
works locally (15 games, 14 with weather, Coors composing to 1.241). So
the failure is downstream of buildContext and has to be located, not
inferred from an absence.
Replays buildContext + contextFor against the CURRENT cached snapshot
grades and counts the drop-off at each hop: team field present -> resolves
to an abbr -> abbr matches a game -> environment produced; and bats /
playerId / opposing-pitcher known -> matchup produced.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
STEP 1 -- input mapped and measured. opportunity_drift 94% coverage on 100
real props: 100% for batters (total_bases, hits, home_runs), 40-67% for
pitchers, which is correct -- pitchers accumulate few at-bats so the ratio
is genuinely undefined and ABSTAINS rather than being invented.
STEP 2 -- THE COLLINEARITY GUARD PASSES DECISIVELY. Pearson r on n=94:
drift vs l20_avg -0.020, vs l5_avg +0.027, vs ab_per_game -0.029. All
essentially zero, so the axis is orthogonal to every existing projection
input and carries information the projection does not already contain.
That also validates the ratio-over-level decision EMPIRICALLY: ab_per_game
is the same quantity over the same denominator as l20_avg, so the level
would have been redundant. Dividing by the player's own baseline removed
the collinearity -- r = -0.029 against the very quantity it is built from.
STEP 3 -- live as a challenger, verified on prod over an induced 416-grade
snapshot: 142 of 276 rows (51.4%) carry the opportunity axis, the
challenger moved on 190 rows, mean |delta| 0.034, range -0.089..+0.108.
Champion p_win and the live grade path are unchanged.
STEP 4 -- HOLDOUT IS n-BLOCKED BY CONSTRUCTION and I am not manufacturing
one. Settled rows carrying the axis: 0. Its first rows carry game_date
2026-08-01 -- games that have not been played. Running the test on rows the
axis never touched would dilute the comparison with rows where challenger
=== champion by construction, making a null result look like a small
positive one. Query committed for when n arrives; it filters to
axis-carrying rows for exactly that reason, buckets before measuring
reliability, and splits time-forward. BOTH metrics must improve or the axis
is shelved.
A MEASUREMENT TRAP RECORDED: the first prod run showed drift at 0% while
ab_per_game read 94% -- indistinguishable from "the feature does not
compute". It was the 120-second feature-vector cache serving payloads
written by the previous image. A new feature field is invisible for one
cache generation after deploy. I nearly reported it absent, having already
confirmed atBats is present in the live statsapi payload and that the code
produced drift = 1.05 locally on that exact data; the contradiction
between those two facts is what saved it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Champion p_win and the live grade path are BYTE-IDENTICAL: the axis writes
only to p_win_challenger / challenger_adjustments in the ledger.
STEP 1 -- MAP THE INPUT. MLB_LOG_FIELD now maps at_bats -> 'atBats'.
Deliberately NOT added to outcomeService's map or liveTracking's
LIVE_BOX_FIELD: those exist to SETTLE and TRACK graded props, and nothing
grades at-bats, so adding it there would imply a settlement path for a
market we do not carry. A test asserts the settle map still lacks it.
STEP 2 -- DRIFT, NOT LEVEL. opportunity_drift = mean(last-5 atBats) /
(season atBats / games). The LEVEL is collinear with l20_avg (same
games denominator; hits/game ~= (hits/AB) x (AB/game)), so the projection
already embeds it multiplicatively and adding it would double-count. A
deviation from the player's own baseline is the part the projection does
not contain.
HONEST ABSENCE throughout: fewer than 3 at-bat rows, no at-bats in the
logs, or no season baseline all leave drift UNDEFINED -- never 1.0 by
default and never 0. Number(null) === 0 here would read as "zero at-bats",
the strongest possible fade, invented from missing data. Four tests cover
the absent paths.
STEP 3 -- THE AXIS. opportunityNudge composes in the same log-odds space
as park and platoon (log of a ratio), with two guards the measured axes do
not need: a +/-10% DEADBAND (a rest day or a blowout can move a 5-game
window without any role change) and a tighter cap (0.15 vs the
environment's 0.30) so a noisy PROXY cannot outvote measured signals.
Every adjustment carries is_proxy: true and
proxy_for: 'confirmed_batting_order' so nothing downstream can mistake it
for a lineup feed.
The axis can stand ALONE -- without it the early return would gate
opportunity off on exactly the thin-classification rows it is most likely
to help.
Zero extra I/O: analyzeViaEngine1 attaches drift from the feature vector
it has already built, and attachChallenger reads it off the grade. Nothing
re-fetches in a loop that runs over hundreds of props.
COLLINEARITY GUARD added to the coverage probe: Pearson r of drift against
l20_avg / l5_avg / ab_per_game, returning null under n=8 rather than
reporting a correlation on a handful of rows. If drift just re-encodes the
projection, the axis is dead signal and gets shelved.
Gates: 4,073 tests / 326 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
READ-ONLY. Live grade path byte-identical -- no layer wired, no threshold
moved, no challenger added, no holdout run.
INPUTS ARE 100% POPULATED (n=80 real MLB props, through the grader's own
path): ab_per_game, rest_days, l5_avg, l20_avg, l10_stddev and
game_count_in_7d all 100%; opp_rank_stat 65% overall and 0% on
stolen_bases. So there is no honest-degradation problem to solve.
FOUR FINDINGS THAT STOP THE WIRING, three of which would have made the
work unmeasurable or wrong:
1. THE PREMISE IS WRONG. There is no built opportunity layer to connect.
ab_per_game is consumed in exactly one place -- analyzeViaEngine1:379,
which renders "4.3 AB/G" on the grade card. engine1 has NO opportunity
or usage factor at all. A projected opportunity was never built;
building one is construction, not connection.
2. THE INPUT IS THE WRONG SHAPE. ab_per_game = season atBats/games. It is
a per-player CONSTANT (measured: varies for 3 of 20 players, and those
cannot be legitimate since the value can't depend on stat_type), so it
can only move all of a player's props together, never separate them.
And it is collinear with the projection: l20_avg = seasonTotal/games,
the SAME denominator, so l20_avg already embeds opportunity
multiplicatively. Adding it additively double-counts.
3. THE REAL INPUT DOES NOT EXIST. depthChartService returns battingOrder:
null for MLB ("the one lineup slot the free schedule feed exposes") and
PropLine /context carries lineup_confirmed as a BOOLEAN, not the order.
4. ARCHITECTURE: wiring it into engine1 would be unmeasurable BY THIS
ORDER'S OWN TEST. Step 2 proves reliability and resolution, both
measured on p_win. engine1 factors move the grade LETTER and never
touch p_win. The layer belongs in probabilityEstimator, which already
adjusts on opp_rank_stat, home_away and a consistency pull.
SEQUENCING IS ALSO STALE: challengerProjection (arch-v1) is already live
with archetype, matchup (platoon) and environment (park) axes, writing
p_win_challenger to the ledger. Step 2 of the order's sequence is partly
done -- and the harness this order needed already exists.
RECOMMENDED INSTEAD, as its own order: an `opportunity` axis on that
harness driven by DRIFT, not level -- recent AB/G (last 5) over season
AB/G. A deviation is not collinear the way the level is. Per-game atBats
is present in the statsapi log rows but MLB_LOG_FIELD never maps it, so it
is a small contained BUILD, which is why it gets its own order. Honest
caveat carried forward: it is still a proxy, not tonight's opportunity.
PROBE BUG RECORDED: the first run reported 0% for every feature including
l5_avg, on a pipeline that had just graded 365 props -- impossible, so the
probe was wrong. getFeatures takes camelCase and returns { features: {} };
I passed snake_case and read the top level. Fixed to call
computeFeaturesForProp. Same class as the earlier silent-false harness: a
measurement that makes working code look broken invites you to "fix"
something that was never broken.
Gates: 4,059 tests / 325 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The first run reported 0% coverage for EVERY feature including l5_avg --
which projectionFor requires, on a pipeline that had just graded 365
props. That is impossible, so the probe was wrong, not the pipeline.
Two bugs, both in my probe: featureCache.getFeatures takes camelCase
(playerName/statType) and I passed the prop's snake_case shape, and it
returns { features: {...} } while I read the top level. Either alone
yields all-zeros.
Now calls computeFeaturesForProp -- the grader's own entry point -- so it
measures what the grade path actually sees. Same class as the earlier
harness that returned a silent false: a measurement that makes working
code look broken is more dangerous than no measurement.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Before wiring any layer into the grade, measure whether its inputs are
actually populated on real props. A layer wired onto sparse inputs does not
degrade gracefully by default -- Number(null) === 0 turns a missing
opportunity into 'zero opportunity', a fabricated input rather than an
absent one.
Reports population per feature, SPLIT BY stat_type, because a feature can
be 100% present for batters and 0% for pitchers and a pooled number would
hide exactly that. Also reports whether ab_per_game varies across a
player's own props -- a per-player constant can only move all of a
player's props together, which is a very different thing from a per-prop
opportunity signal.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
runSnapshot now takes an opts object, so the route call is ('mlb', {}).
Asserted as EMPTY rather than loosened to any-object: a stray limit
reaching production would silently cap every run, which is the exact bug
the hook exists to diagnose.
I pushed the previous commit without reading the suite result -- the
failure was already on screen. Caught and fixed immediately after.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Induced, not projected. DEFAULT_LIMIT=500 produced 365 graded props in
114s (was 7 in 16s) -- 52x the board. All 365 carry a unique forecast_rank
and ZERO leak p_win to anonymous callers, so the tier gating holds at 50x
the volume. Anon payload 220KB in 0.44s. Stat mix went from three stats to
ten. Health green.
Measured cost curve via the ?limit= bisect hook: 1->42s, 25->58s, 60->42s,
120->66s, 500->114s. About 42s of that is FIXED overhead (odds fetch,
roster logs, archetype classify, retention), paid whether we grade 1 prop
or 500 -- grading is the cheap part.
MY PRE-FLIGHT ESTIMATE WAS WRONG. I predicted ~72s from per-prop latency
measured in isolation, which ignored the fixed cost. Real figure 114s.
A FALSE ALARM RECORDED because acting on it would have meant reverting a
fix that works: the first induced run 502'd at 13.4s and I hypothesised
load -- memory or a proxy timeout under 20x the work. Wrong. A limit=25 run
then 502'd in 2 seconds, which no amount of load explains, and both
recovered on retry. The 502s were the deploy rolling, not the cap.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The cap raise 25 -> 500 made an induced snapshot 502 at 13.4s and the run
did not complete in background either, while a 25-prop run had completed
in 16.3s. That rules out a simple duration timeout and means the cause has
to be measured, not guessed. ?limit= bounds one run so the regression can
be bisected without a prod env change; omitted, the real DEFAULT_LIMIT
applies.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
PART 1 (read-only, measured on a live prod slate, n=80) OVERTURNS THE
PREMISE. The refusal rate is not a data problem -- it is 98% correct
behaviour. The cap is the entire problem, and it is worse than "25 of 546".
Composition: GRADED 44 (55.0%) | POLICY-SUPPRESSION 35 (43.8%) |
FETCHABLE-GAP 1 (1.3%) | FALSE-THRESHOLD 0 | ARCHETYPE-GAP 0 |
GENUINE-ABSENCE 0.
THE FIFTH BUCKET the order did not anticipate: all 35 "refusals" are
rare_event_over_below_line -- the 2026-07-19 betting-logic audit
deliberately refusing 0.5-line rare events, setting the SAME
insufficient_data flag as a real data gap, which is why they read as one.
They are entirely doubles (18) and stolen_bases (17), while hits (19/19),
rbi (19/19) and total_bases (5/5) grade at ~100%. Had we "fixed" this we
would have re-introduced exactly the bets a previous audit removed, and the
count would have looked like progress.
THE CAP: 585 unique gradeable props, cap 25 -> 560 discarded (95.7%).
Traced to Session 32 (f0c8b4f), commented "bound the herd" -- a guard
written before anyone measured what a grade costs. So I measured it:
721ms mean / 666ms median / 1024ms p90 per grade => ~72s for 500 props at
concurrency 5. Both callers tolerate that: the cron runs 5x/day and
recordDownstream is fire-and-forget.
PART 2 -- item 3 ONLY, because that is what the diagnosis supports.
DEFAULT_LIMIT 25 -> 500, env-tunable via GRADE_SLATE_LIMIT. Concurrency
stays 5 deliberately: the cap raise already multiplies load ~20x, and
concurrency decides how hard we hit statsapi at once. One variable at a
time.
Items 4/5/6 have nothing to act on and I am not manufacturing work for
them: 0 false thresholds to loosen (loosening would be manufacturing
grades); /context wiring is worth doing for grade QUALITY but would not
have graded one extra prop here, so it is not claimed as a coverage win;
archetypes are display-side and do not gate grading at all.
THE REFUSAL RATE DOES NOT DROP, AND THAT IS CORRECT. No threshold lowered,
no grade forced. The board grows because the cap stops discarding 95.7% of
the slate.
Flagged in advance rather than discovered later: snapshot payload and
ledger volume both scale with the same multiple. If the response gets
unwieldy the fix is a response-side cap on what the BOARD returns, never a
re-cap on what gets graded -- grading everything and serving a slice is
honest; grading a slice and calling it the slate is what this fixes.
Gates: 4,052 tests / 324 suites green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
READ-ONLY. Runs the REAL grade path over a REAL slate and categorises
every refusal; writes nothing. Reproduces gradeSlateService.dedupeProps
exactly (MODEL_BOOKS, first-row-wins) and calls analyzeViaEngine1 the same
way, so it measures what the pipeline does rather than a re-implementation.
Adds a FIFTH bucket the order did not anticipate, and it is likely to
change how the 72% is read: (e) POLICY-SUPPRESSION. The 2026-07-19
betting-logic audit deliberately refuses rare-event 0.5 markets (doubles/
triples/HR/SB) on the juiced under, plus any over-juiced price -- and it
sets the SAME insufficient_data flag as a genuine data gap. Counting those
as a data problem would send us hunting for data that is not missing, and
"fixing" them would re-introduce bets we removed on purpose.
Separates (b) FETCHABLE-GAP from (d) GENUINE-ABSENCE by asking the stats
layer directly whether the player has ANY game log, rather than assuming:
no log -> genuine absence, keep refusing; a log that exists while the grade
path found no projection -> a wiring gap with something to fix.
Also measures per-grade latency (mean/median/p90/max, serial and at
concurrency) so Part 2 can decide the cap on cost rather than on taste.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
PART A -- WNBA TRUTH CORRECTION (no behaviour change).
WNBA does not "abstain" and is not "anti-predictive". The -0.12 that
produced those words was NBA-template machinery run on WNBA data -- WNBA
has never had its own archetypes, variables, conditions or calibration,
which is precisely the "sport stubbed in on another sport's template"
CLAUDE.md forbids. That is an UNBUILT MODEL'S EXPECTED FAILURE, not a
verdict on the sport; reading it as a verdict would quietly retire a sport
we never actually attempted. Its own build is QUEUED, after MLB.
The guard CODE is unchanged -- FORECAST_RANKED_SPORTS = {'mlb'} and the
inheritance test are correct live safety either way. Only the meaning is
corrected, and generalised into the doctrine-as-a-gate: a sport ranks on
p_win ONLY once its OWN model is built and shown to predict (calibration
AND resolution on its own holdout). Others are held out as NOT-BUILT,
never as failed. Re-labelled across gradeRanking, snapshot route, tests,
MASTER-PLAN and the challenger report.
PART B -- THE FLIP, gated on a full-slate re-run.
The re-run found something better than a bigger sample. An induced
snapshot graded 7 props: gradeAndCacheSlate runs with DEFAULT_LIMIT = 25
and ~72% of those refuse for insufficient_data, while 546 props are
gradeable. So 8 props IS the board, structurally -- not a small sample of
it. Logged as its own finding; the cap is a separate order.
For a statistically meaningful delta I used 11 real historical boards
(n=328, board sizes 14-57): 79.9% of rows move, mean 5.16 places per
board, TOP READ CHANGES ON 9 OF 11 BOARDS. The re-ordering holds at real
board size. Query committed.
FLIPPED:
- rankGrades drops its edge key (safe for every sport: removes a
non-predictive tiebreak without putting p_win in front).
- selectTopGrades leads on forecast_rank, edge key removed.
- flattenToEdgeBoard sorts on forecastRank, not edge -- this board had
edge as its PRIMARY key, so the whole mobile board was ordered by a
quantity measured not to predict.
- forecast_rank threaded onto strip props.
Sports whose model is not built supply no forecast_rank, so their boards
fall through to the unchanged grade chain -- the fallback is the guard.
ROLLBACK ARMED: boards sort by forecast_rank WHEN PRESENT, so
FORECAST_RANK=0 reverts every surface on the next response -- no deploy,
no client release.
Edge is still computed, stored, carried and displayed as a labelled
diagnostic. Retired from ranking, not deleted.
Eight superseded tests updated to strictly stronger INVERSE properties --
they now fail if edge is ever re-introduced as a ranking key, which the
originals could not detect.
Gates: 4,045 tests / 323 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
DELTA MEASURED on live prod grades (live ordering unchanged): MLB 7/8
props move (87.5%), mean 2.5 places, TOP READ CHANGES (corey seager hits
1.5 under -> jake burger hits 0.5 over). WNBA 25/25 move, mean 4.1, max 12.
This is a large re-ordering, not a tweak.
Caveat recorded rather than buried: MLB had only 8 graded props at
measurement time. The percentages are real; the sample is one small slate.
Re-run before the flip -- it is one call.
PER-SPORT DOCTRINE ENFORCED IN CODE. WNBA moves the most and must NOT
adopt this: its p_win is anti-predictive, so ranking that board by p_win
would sort it by a signal measured to point the WRONG WAY -- worse than
the incumbent, not better. A comment would not have stopped a future flip
from going global, so FORECAST_RANKED_SPORTS = Set(['mlb']) gates the
forecast_rank stamp, with tests asserting no sport inherits MLB's result.
A sport joins only by passing its own holdout.
EDGE IS NOW DIAGNOSTIC-ONLY IN DISPLAY. MobileEdgeBoard.EdgeCell rendered
green (--g-a) for positive edge and red (--miss) for negative. Two things
were wrong: green/red IS a quality claim on a quantity that does not
predict, and ROW-GRAMMAR reserves red for settled-negative ONLY -- a
negative diagnostic is not a settled loss. Now neutral mono with a
diagnostic tooltip; header reads "MKT GAP · DIAGNOSTIC". The number is
still shown -- no display went blank. DeskShowcase neutralised likewise.
PINNACLE LOGGED, NOT ENSHRINED. Per the order, "market-not-sharp" is
PENDING-RECOVERY rather than a confirmed permanent limitation. The single
question for PropLine is in BLOCKERS.md with its evidence, and MASTER-PLAN
now carries the pending status instead of the permanent claim.
Live sorts remain byte-identical: selectTopGrades, flattenToEdgeBoard and
topGradedService all still call the incumbent.
Gates: 4,041 tests / 323 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
MEASURED BASIS (n=200 settled MLB rows): corr(p_win, outcome) = +0.26;
corr(edge, outcome) = -0.010 incumbent ruler / -0.022 consensus ruler.
Subtracting the market destroys the signal under BOTH rulers, so a
quantity that does not predict must not rank, gate or decide.
CHALLENGER-FIRST -- live ordering is byte-identical. rankGrades (the
incumbent, grade-first with edge as its 4th key) is untouched and tested
as untouched.
NEW: rankByForecast -- takeable-gated p_win -> grade -> confidence -> stable
order, with NO edge term anywhere. p_win LEADS and the letter follows,
deliberately: the letter measured r ~ 0.005 and is inverted (B 52.4% <
C 56.9%) while p_win measures +0.26, so leading with the letter would sort
by the weaker signal and use the stronger one only to break ties.
Recorded in the code: isotonic calibration is a MONOTONE transform, so
ranking on raw vs calibrated p_win gives the SAME ORDER. Calibration
matters when p_win is displayed or thresholded; it cannot change a
ranking. Nothing here needs the calibrated value.
rankingDelta + GET /api/internal/ranking-delta measure how far the board
would move before any flip. The endpoint reports p_win coverage alongside
the delta -- if p_win is absent the challenger degrades to grade order and
the delta UNDERSTATES, which is worth saying rather than reporting a clean
zero.
forecast_rank is stamped on snapshot grades BEFORE stripModelPrice, so
every tier gets the correct order without the paid values (the
topGradedService precedent -- an ordinal can travel where the magnitude
cannot). Additive only: nothing sorts by it yet.
RETIRED AS DECISIONS (not rankings, so done now):
- altLineScanner.compareToBookImplied no longer returns value_detected:
edge > 0. Edge is still COMPUTED and returned -- losing the record would
be worse than mis-using it -- but the verdict is an honest null with
value_basis: 'retired:edge_does_not_predict'.
- scanAltLines no longer filters to edge>0 or calls the survivor "optimal".
The whole ladder is returned ranked and labelled
'price_gap_diagnostic_unvalidated'. The module has ZERO callers (verified)
-- unwired like mlbGrader.js, left in place and made honest.
An honest asymmetry recorded there: ranking props AGAINST EACH OTHER must
not use edge, but choosing between RUNGS OF THE SAME PROP is inherently
price-relative -- ranking rungs by model probability alone would always
pick the lowest line, since P(over 0.5) > P(over 2.5) by construction. So
the gap stays the rung key, explicitly labelled unvalidated.
Two superseded tests updated to stronger properties.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
MEASURE-ONLY. No promotion, no flip, no tier spend. Live path
byte-identical: CURRENT_RULER_VERSION still v1_first_book, model still
consumes MODEL_BOOKS only.
MANDATE 1'S PREMISE DOES NOT HOLD. The p_win calibration is
RULER-INDEPENDENT, confirmed two ways: estimateProbability takes
{gameLogs, line, statType, features} and never sees a market price, and
the calibration fits p_win against OUTCOMES. Reliability and resolution
are both p_win-vs-outcome measures, so fair_prob cannot enter either.
There is nothing to re-fit -- the ruler changes edge, CLV and takeable,
not calibration.
I RETRACT MY OWN LABEL. I declared the MLB isotonic result PROVISIONAL
"because it was measured against the bent ruler". That over-applied the
ruler caveat to a measurement the ruler never touched. The result was
never contaminated; it moves PROVISIONAL -> DECIDED, not by re-running but
because the gate I attached does not apply.
RAN THE GENUINELY RULER-DEPENDENT QUESTION INSTEAD -- does a median
consensus rescue EDGE? Timing held constant (both rulers at close; a
lock-time reconstruction joins only 43 rows, and mixing lock-incumbent
with close-consensus would confound WHEN with WHAT).
n=200 MLB settled rows: mean |ruler gap| 0.0085. corr(edge_v1, outcome)
-0.0101; corr(edge_v2, outcome) -0.0220; corr(p_win, outcome) +0.2598.
THE HEADLINE: p_win predicts outcomes at +0.26 while p_win minus the
market predicts nothing under EITHER ruler. Subtracting the market price
destroys the signal -- a direct empirical vindication of the identity now
at the top of CLAUDE.md. Market edge is not merely a poor criterion here;
it is a strictly worse instrument than the raw forecast.
CALIBRATION REFRESH (ruler-independent, but n grew 119 -> 250):
time-forward holdout n=125, reliability 0.0846 (was 0.0939), resolution
0.190 (was 0.123). Both hold and both improved on a fresh later window
the earlier fit never saw. Independent replication.
THE LIMITATION THAT BLOCKS A FULL VERDICT: closing_captures holds only
MODEL books -- exchange quotes were never stored, because normalizeProps
discarded them until yesterday. Mean 1.97 books in the historical join. So
this tested a US-books-median ruler, not the exchange-inclusive consensus
whose live delta showed p90 +10 points. That ruler is UNTESTABLE on
existing data at any n. Per Mandate 4's third outcome: inconclusive, not
forced.
SEPARATE FINDING -- LIVE FEED REGRESSION: pinnacle MLB captures went 4,022
-> 0 on 2026-07-31 and have not returned, while every other book continued
(103,940 captures in the prior 10 days). This also corrects an Order Zero
claim of mine: "no sharp anchor exists in our feed" was accurate for the
day measured but wrong generally -- pinnacle was there until 07-30 with
17,090 two-sided captures. line_type='sharp' is a label in closingCapture
via SHARP_BOOKS, not a separate provider. We had a sharp anchor and lost
it two days ago; not caused by anything in this session.
Both queries committed: scripts/ruler-comparison.sql,
scripts/pwin-timeforward.sql.
Gates: 4,028 tests / 322 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
PHASE 1 resolved on our real keys, and a bad source was discarded on the
way: a fetched rendering of PropLine's docs "tier matrix" claimed
/odds/closing is 403 on free and that /odds returns prices nulled on free.
Both are contradicted by direct observation (200-redacted, and 6,196
two-sided PRICED groups on MLB). Not cited. The report uses only the
machine-readable OpenAPI contract and the verbatim detail bodies our keys
received.
Verdict: every one of the six endpoints behaves exactly as the Free tier's
published contract says. error:"upgrade_required" with an explicit
required_tier is unambiguous -- NOT a key-permission problem, NOT a plan
problem. $9/mo Hobby buys /results + /odds/closing (the CLV instrument) +
/movement (steam across 18 books); $19/mo Pro adds the 90-day settlement
export. Priced and evidenced; not recommended here -- it is a decision.
PHASE 2 fingerprint on the SERVED feed: 5 books -> 13, props rendered
546 -> 2,780 (5.1x), mean 4.22 books/prop.
The unflattering half, stated up front: of 2,234 newly-visible props only
698 (31.2%) carry a real non-DFS market price; 1,536 (68.8%) are DFS-only
pick'em rows. The honest headline is not "80% of the slate unlocked" --
the board is 5x fuller, about a third of the new depth is real market
data, and the rest is pick'em inventory now shown but tagged.
PHASE 3 verified: 546 gradeable props, unchanged. CURRENT_RULER_VERSION
still v1_first_book.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The route regroups flat props into lines[] and was dropping the role tag,
so the widened feed reached the browser untagged. That is not cosmetic:
on a live prop, PrizePicks prices both sides at even money (+100/+100)
while BetMGM has +450/-750. Rendered side by side without a tag, the
pick'em row reads as a dramatically better price when it is a different
product entirely -- exactly the confusion the three-way split exists to
prevent. Consumers gate on book_role !== 'dfs' before treating a row as a
market price.
The ?book= filter now accepts any DISPLAY book, since shopping a real
book against an exchange is the point of the widening. Grading still only
ever consumes MODEL_BOOKS.
One superseded integration test updated to a stronger pair: an unknown
book still 400s, and a newly-visible one no longer does.
Gates: 4,028 tests / 322 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
IDENTITY (CLAUDE.md top + MASTER-PLAN header). VYNDR is a PREDICTIVE MODEL:
it projects what a player will DO and picks accurately. Market edge is a
BYPRODUCT of a good prediction, never the success criterion. Success =
the forecast is honest about its own confidence AND still ranks --
calibration and resolution, both. No edge/CLV term belongs in a pass/fail
gate; they are diagnostics we report, not thresholds a model must clear.
A model tuned to beat a closing line has been fitted to the market instead
of to the game.
Per-sport doctrine (Phillips 2022, classify by what players DO not by
position): each sport is its own model -- own variables, archetypes,
conditions, calibration, honest ceiling. Shared across sports: ONLY the
Bayesian inference math.
Truth Law: no fabricated data; honest-absent over invented; label
limitations in-band; provisional stays provisional until re-run;
documented is not verified.
PHASE 2 -- AGGREGATOR WIDENING (live). normalizeProps now emits every
DISPLAY book instead of 5 of 18. Before this we discarded 13 books of our
own accord and 64.8% of the MLB slate was invisible to users. Every prop
carries book_role (both/takeable/reference/dfs/offshore) so the display
layer can say WHAT a price is -- a fixed-payout DFS number and a two-way
sportsbook price are not interchangeable objects. Unknown books are still
dropped.
PHASE 3 -- MODEL GATE (the model does not move). bookRoles splits
MODEL_BOOKS (the legacy allow-list, character for character) from
DISPLAY_BOOKS. Both model paths re-filter before they pick a line:
gradeSlateService.dedupeProps (before first-row-wins AND before the limit)
and intradayRefreshService.indexOddsProps (which RE-GRADES at the current
line -- without the gate, widening would have silently moved locked lines
onto books the model has never been calibrated against). A test asserts
the graded set is byte-identical through the widening.
CURRENT_RULER_VERSION stays v1_first_book. The gate lifts only when the
MLB calibration is re-run on the consensus ruler and v2 is promoted.
HONEST FRAMING, recorded in the plan: this is an AGGREGATOR win and it
does NOT fix the model. WNBA still abstains -- a model problem, not a
coverage problem; it is better covered than MLB. MLB isotonic still
provisional. The consensus is MARKET, not SHARP: pinnacle, matchbook and
polymarket are 0% on both sports, so no sharp anchor exists in our feed.
Two superseded tests updated to stronger properties rather than deleted:
roleOf now names the KIND of book, and the normalizer test asserts the
display set widens WHILE the model set does not.
Gates: 4,027 tests / 322 suites green; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
PHASE 1 (measured on the live prod feed with the real key):
- WNBA is NOT thin at the feed -- 4.21 books/prop vs MLB's 3.61. It was
allow-list-starved exactly as MLB was. This removes one candidate
explanation for its anti-predictive result; it does not explain it, and
WNBA stays abstaining.
- We cannot see 64.8% of the MLB slate at all (zero admitted books).
- Exchanges are real (smarkets 27%, novig 22%, kalshi 15% on MLB) but
pinnacle, matchbook and polymarket measured 0% on BOTH sports. There is
no sharp anchor for player props. The consensus is a MARKET consensus,
not a SHARP one -- recorded as a permanent limitation, not a milestone.
- DFS is the trap, quantified: prizepicks covers 82% of MLB props, the
highest in the feed. Admitting it "for breadth" would have looked like
the biggest available win. Permanently excluded.
- Endpoints: /context WORKS and is FREE (umpire, roof, pitcher handedness,
lineup confirmation -- richer than what we hand-built). /odds/closing and
/movement are REDACTED (full structure, zero prices). /results and
/exports/resolved-props are 403.
- The $19/mo question is answered: soccer IS graded, ~15 competitions in 30
days (MLS 41k, Liga MX 15k, Brasileirao 12k, UCL/Europa/Conference). Our
"soccer grades into a void" is a Pro-tier problem, not a data problem.
NBA is absent because it is July -- seasonal, not inferable either way.
PHASE 2 delta, corrected: MLB mean +1.50 pts, median 0, p90 +10.0, 17.0%
of comparable props move >=5 pts, one-directional (the incumbent prices
the over below the exchange-inclusive consensus). WNBA symmetric and
tight. The median prop does not move -- the change is a right-skewed
minority. That the rulers DIFFER is established; that the new one is
BETTER is not, and that is the re-run.
PHASE 2 item 6: ledger_entries.ruler_version applied to prod, 1,384
existing rows backfilled to v1_first_book (a statement of fact -- every
row to date was produced by the first-book rule). ledgerService stamps
CURRENT_RULER_VERSION on new rows. Never pool edge or CLV across it.
Repo migration numbering lags prod; 025_ledger_ruler_version.sql records
the DDL for review.
PHASE 3: MLB isotonic p_win remains PROVISIONAL -- calibrated against
v1_first_book, does not promote until re-run on the consensus ruler.
NOT LIVE, deliberately: ALLOWED_BOOKS unchanged, served slate
byte-identical, CURRENT_RULER_VERSION still v1_first_book, no live path
calls consensusRuler.
Gates: 4,022 tests passed / 322 suites; next build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
My first delta run modelled the incumbent as first-row-wins over the RAW
feed and reported that an EXCLUDED book was "the market" on 69% of MLB
prop-lines, with prizepicks alone at 47%. That is WRONG and I caught it
before it went anywhere.
normalizeProps applies ALLOWED_BOOKS BEFORE gradeSlateService.dedupeProps
runs, so DFS books never reach the incumbent. The allow-list, for all the
coverage it costs, does keep DFS out of the ruler.
incumbentFairProb now takes the allow-list (defaulting to the live
ALLOWED_BOOKS) and reproduces the real chain. Two tests lock it, including
that a prop with no admitted book has NO incumbent -- it is never graded
at all, which is the real loss and is already measured as invisible_props.
Overstating the incumbent's badness would have been as dishonest as
understating it, and more persuasive.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
CHALLENGER-FIRST. The live ruler is byte-identical: CURRENT_RULER_VERSION
is still v1_first_book, nothing here writes a cache, a grade or a ledger
row, and no live code path calls consensusRuler yet.
bookRoles.js splits one allow-list into three, because it was answering
two different questions -- "can we show this?" and "can we price against
this?" -- with the same list, which is what bent the ruler.
TAKEABLE the user can actually bet here (drives best price / shopping)
REFERENCE may price the fair-prob ruler; never surfaced as a place to bet
EXCLUDED DFS pick'em + offshore, permanently barred from all pricing
Two deliberate calls, both evidence-based:
- The six PropLine-phantom books (caesars/fanatics/bet365/hardrockbet/
pointsbet/thescore) are KEPT despite the order saying remove. They
returned zero PropLine quotes, but PropLine is not our only provider and
the odds-api backup path may carry them. A book that never appears is
never matched, which costs nothing; deleting them risks silently
dropping real books on the backup with no upside. Recorded in
PHANTOM_ON_PROPLINE rather than enacted as a deletion.
- REFERENCE = exchanges + pinnacle + bovada + the four US majors, chosen
off the measured coverage curve rather than theory. exchange_only is
cleanest (order-book, ~zero vig) but covers 14.3% of MLB and 5.6% of
WNBA; adding the US majors gives 28.1% / 46.3%. pinnacle, matchbook and
polymarket measured 0% on both sports and add nothing. The honest
limitation is recorded in the config: this is a MARKET consensus, not a
SHARP one.
consensusRuler.js: median de-vigged fair_prob across >=2 reference books
posting BOTH sides at the SAME line. Median so one stale exchange cannot
drag it. Different lines are never averaged, one-sided quotes never rule,
and n<2 falls back to single-book LABELLED as such with the v1 stamp --
never silently mixed, because a column holding both is two rulers wearing
one name.
The challenger delta runs over the live feed and reports incumbent_book_
roles, which is the real headline: the incumbent is literally first-row-
wins, so it reports what KIND of book has been acting as "the market".
DFS pick'em has the highest coverage in the feed, so a DFS book can be it.
18 ruler tests + 37 total in the two new suites. Full suite 4021 passed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
/sports and /markets/resolution-summary carry no per-prop data and no
credentials, and the shape summary alone cannot answer the question they
exist to answer -- whether PropLine actually GRADES the sports we cannot
settle. A shape is not a number. Both bodies are scrubbed on the way out.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Two corrections to the first pass, both of which would have produced a
false positive.
1) A non-empty body is NOT proof of access. PropLine's free tier returns
the full STRUCTURE of tier-gated endpoints with values stripped plus an
upgrade_url -- and the first pass classified /odds/closing and /movement
as "works" on structure alone. detectRedaction() now counts actual
prices and downgrades works -> partial when a body advertises an upgrade
or carries outcomes with zero prices. Same class as the harness that
returned a silent false, inverted.
2) One hard-coded reference set forces a yes/no on a question that is
really a curve. reference_policy_curve reports strict eligibility
(>=2 books, both sides, same line) under exchange_only /
exchange_plus_sharp / exchange_plus_us / takeable_only, so the ruler
decision is made on coverage-vs-quality rather than on a guess. DFS is
absent from every policy by construction and a test asserts it.
Also probes /markets/resolution-summary: /exports/resolved-props being 403
tells us we cannot PULL settlements; resolution-summary tells us whether
they EXIST to be bought. Different questions.
19 unit tests, still hermetic.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Adds GET /api/internal/propline-verify (internal-key gated, read-only) so
Phase 1 can run WHERE THE KEY LIVES. Touches no cache, no ledger, no
grade; the live adapter and the live ruler are untouched. Breadth reuses
proplineAdapter.fetchRaw -- the exact live request -- so what it measures
is what the pipeline actually receives.
Reports per sport (never pooled): books/prop from the feed vs after our
own ALLOWED_BOOKS, props made INVISIBLE by that filter, reference-book
presence, DFS presence reported separately, and consensus eligibility.
Consensus eligibility is deliberately strict: >=2 REFERENCE books posting
BOTH sides at the SAME line. A one-sided quote cannot be de-vigged, and
two books at different lines are not the same market -- counting either
would overstate how much of the slate can carry a real ruler.
Probes the documented-but-unverified endpoints (/sports, /context,
/odds/closing, /movement, /results, /exports/resolved-props for four sport
keys) and classifies works/partial/no, with 403 = tier-gated and 200-but-
empty = partial rather than works.
Key safety is the other locked property: the key goes via axios params,
never string-interpolated, and every emitted string passes scrubKeys()
which removes the literal key AND any surviving apiKey= query value. A
test asserts a thrown transport error carrying the key cannot escape.
13 unit tests, hermetic (no network, no key).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
STEP 0 disproved the premise before any request was fired. PropLine's
OpenAPI contract states verbatim that `bookmakers` omitted = ALL books,
so proplineAdapter omitting it is correct and always was. Firing a
guessed param would have RESTRICTED the response and produced exactly
the false negative the order warned about.
The real cause is ours: PropLine sends 18 books; oddsNormalizer
ALLOWED_BOOKS intersects them at exactly 5 -- which is precisely the
"5 MLB books" the 2.18 audit measured. Measured on real public data
(no key, no quota): 4.41 books/prop from the feed, 1.50 after our
filter, and 12 of 34 props go invisible entirely.
Also corrected: "73% single-book" is the long tail of deep props
sole-posted by DraftKings or Bovada. On the core props we grade, the
market is 10-12 books wide. pinnacle appears on 0 of 40 MLB props --
the independent low-vig references present on 100% of core props are
exchanges (novig/smarkets/kalshi). DFS pick'em also covers 100% but is
not a market price and must never enter a consensus.
Verdict is outcome (d) ALREADY OPEN, not (a)/(b)/(c) -- all three
assumed the feed was the constraint. Ruler change scoped (not built):
split one allow-list into takeable/reference/excluded, fair_prob_lock
becomes a median consensus with n>=2 or a labelled fallback. Gated on
exchange price validation + the WNBA measurement, which needs the
PropLine key (prod-only, absent locally). MLB isotonic p_win declared
PROVISIONAL until re-run on the real ruler.
Side finding: we use 1 of 29 endpoints. /odds/closing, /movement,
/odds/history, /best-line, /ev, /results, /exports/resolved-props,
/context (free) map directly onto documented gaps -- and resolution
across 33 sports suggests "no free settled feed for NBA/soccer" may be
a $19/mo problem, not a data problem. Documented, not verified.
Plan edits: §10.1 rewritten, §10.2/§10.5 corrected, and §11 adds the
sequential post-completion accrual clock -- pre-completion data does
not count, no pooling across the completion boundary, two clocks
stated separately, per-sport clocks, verification gate before any
accrual, users onboarded to a complete product only. §9.1's "6-10
weeks out" corrected: that is accrual duration, not distance to the
answer. The ruler change independently forces the same no-pooling
boundary by arithmetic.
No API key was used, printed, or committed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Answers "what makes this the top product, not just a finished one."
THE REFRAME: the aggregator gap and the model gap are the SAME gap in two places.
Our "market" is often ONE book — MLB props are 73% single-book, and
proplineAdapter sends only {apiKey, markets} with NO regions/bookmakers param
(:152), so we take PropLine's default response. That single fact causes four
problems we had been treating as unrelated: no line shopping (the category's #1
free hook), a fair_prob_lock that is a de-vigged single soft book rather than a
consensus (the bent ruler the model is judged against), weak CLV (cannot measure
beat-the-close against one book), and no steam/disagreement detection (needs >=2
books to exist).
So the highest-leverage unblocked action in the whole plan is a cheap API test:
does PropLine return more books with a regions/bookmakers param on our tier? One
request, and if it works it upgrades the free product, the model's denominator and
the CLV instrument simultaneously.
Aggregator gaps catalogued: book breadth, true consensus, historical odds archive
(started — closing_captures 844k rows, lock_lines new, but in-grade history capped
at 24 points, so no full open->close series), market breadth (11 live vs the
category's 50+), ingested alt-line ladders, injury/lineup wire, player news.
Paid-model gaps catalogued: distribution instead of a point (distribution.js
already computes survival probabilities and rungs but is proj-v1.1, ledger-only
and lost to the champion); opportunity/playing-time projected FIRST with its own
uncertainty (the single biggest available modelling gain); per-stat models instead
of one additive index; matchup granularity that actually reaches the grade;
applied calibration; a backtest harness (blocked by the archive gap — you cannot
backtest a price you never stored); CLV as north star.
THE PATTERN: almost every model capability is ALREADY BUILT AND DISCONNECTED.
VYNDR does not have a building problem, it has a connection-and-proof problem plus
one genuine ingestion gap that starves both halves. The expensive part is largely
done, but no new feature fixes it.
Ordering principle recorded: get MLB genuinely good BEFORE replicating across six
sports — a copied-six-times thin model is six times the maintenance for the same
absent edge.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The phases counted unbuilt code. This section names what is missing for VYNDR to
do what it claims, including the parts that are not builds.
THE CENTRAL GAP: there is no demonstrated edge yet. Every measurement this session
returned null, negative or unproven — served grade r~0.005 and inverted; all three
p_win-vs-fair_prob formulations negative on both sports and both splits; p_win
alone on MLB holdout p~0.07; WNBA negative; CLV null by guard; ROI-by-grade likely
an artifact. The product's core claim is not currently supported by our own data,
and building all 23 orders without closing this leaves a well-built product that
does not do the thing it sells. What closes it is sample and honest iteration, not
code — roughly 6-10 weeks at the current accrual, a clock engineering cannot
shorten and that must not be faked.
Also named: the projection is thin (l5/l20 + opponent rank + rest + usage, with
similarity/archetypes/conditions/Bayesian all built and disconnected, so
connecting them is a hypothesis not a guarantee); it is a one-sport product today
(NBA and soccer do not even settle); there are 3 users and 0 paid so nothing is
validated by usage; there is NO distribution path at all, which appears in no
phase and belongs on the board as its own track; the last mile is unclosed
(push-to-book is a teaser, no affiliate live); and operational fragility remains
(single box, two-sport settlement, three credentials flagged including a Stripe
live key that transited a transcript, no staging).
The honest summary: the truth infrastructure is genuinely well built and this
codebase does not lie about what it knows. What is not yet true is that the model
beats the market — not disproven, unmeasured at adequate n. The finish line is 23
orders PLUS a verdict from accrued data we cannot rush, and the discipline to
report that verdict honestly if it says the edge is not there.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Consolidation only. Nothing built, wired or promoted.
NOTHING WAS RE-VERIFIED and no query was run — all 22 artifacts produced this
session plus the completion matrix were taken as KNOWN, per the order's own clause.
The verification ledger at the top of the plan lists exactly what was taken as
known and which four items remain genuinely open (sport order, board-reasoning
gating, the CLV flag, team colours) — each open because it needs a decision or a
build, not a query.
The plan captures all six tracks in one document: per-sport models (MLB's 8-layer
stack with each layer marked BUILT/PARTIAL/NOT-WIRED, plus the sport order),
design implementation (61 catalogued items), surfaces, the resolution tail, the
sport boundary, and the Chrome audit.
The through-line it makes visible: MLB's layers 2, 3, 5 and 6 are BUILT AND NOT
CONNECTED, while layer 8 (the grade ladder) is connected and meaningless
(r~0.005, inverted). MLB's fix is connection, not construction.
Phasing is by dependency: MLB model truth -> resolution tail -> surfaces/design
(parallel lane) -> sport boundary -> sport rollout (one order per sport) ->
monetization finish -> Chrome audit and hardening. ~23 orders total, ~11 unblocked
today, so "how many sessions left" now has a real answer.
DEFINITION OF DONE is explicit and countable: MLB layers 1-8 connected with a
monotone held-out-proven ladder; every listed sport finished on the same template
or explicitly abstaining with its reason recorded; all 61 design items built; every
surface reachable and honest; the resolution pipeline firing end-to-end; the sport
boundary a registry; the Chrome audit passed; and the record publishable on its own
terms with no claim outrunning its evidence.
STATE.md now points at the plan and is demoted to history.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Measure-only. p_win not flipped live, no grade rebuilt, no calibrator deployed.
Per the doctrine, MLB and WNBA were fitted, selected and judged as SEPARATE
models — and they reach opposite verdicts. No global instrument was fitted.
METHOD: time-forward split per sport (earlier fits, later proves). Both
instruments fitted on TRAIN only — single-parameter Platt and
isotonic-with-pooling. Inputs p_win + outcome only; no market field, no closing
value, no lookahead. Nothing about edge/CLV/beat-the-close enters any pass/fail
line.
MEASUREMENT CORRECTION made mid-run: the first pass reported mean|p - outcome|
(~0.46-0.51), which is NOT calibration — it is noise-dominated individual error
on 0/1 rows and would have made every instrument look identical. Reliability is
only meaningful on BUCKETS (bucket mean predicted vs bucket actual rate,
n-weighted), the metric T0 used. All reported numbers use the corrected metric.
HOLDOUT RELIABILITY (lower better): MLB n=119/4 buckets — raw 0.1038, Platt
0.1120, ISOTONIC 0.0939. WNBA n=93/3 buckets — raw 0.1322, Platt 0.0491,
isotonic 0.0667.
HOLDOUT RESOLUTION: MLB raw 0.1388 -> Platt 0.1284 -> isotonic 0.1225.
WNBA raw -0.1201 -> Platt +0.1269 -> isotonic +0.0322.
Fitted Platt: MLB a=-0.381 b=+0.705; WNBA a=+0.040 b=-0.081.
MLB QUALIFIES, MODESTLY — instrument selected BY HOLDOUT, not assumed: isotonic
beats both raw and Platt, and Platt actually made MLB worse. Reliability improves
0.1038 -> 0.0939 (~10% relative, real but modest) and resolution SURVIVES
(0.1388 -> 0.1225, not crushed). Both Mandate-3 conditions hold.
WNBA ABSTAINS — its Platt result is the best number in the report and is REJECTED
as a fake win. The fitted slope is b = -0.081, negative and near zero, so
sigmoid(0.040 - 0.081*logit p) is nearly constant at ~0.51 for every input: it
"calibrates" by discarding the prediction and emitting the base rate, which is
exactly the failure Mandate 3 pre-registered. Its apparent resolution gain
(-0.120 -> +0.127) is the sign flip, not skill — it would serve the opposite of
its own forecast, fitted on n~96 of anti-signal. Isotonic says the same quietly
(resolution collapses to +0.032).
HONEST CEILING: MLB is a usable-but-unimpressive forecaster (holdout resolution
~0.12, reliability ~0.094, n=119); WNBA has no honest forecast today. Holdout n
and bucket counts (4 and 3) suffice to reject WNBA and prefer isotonic for MLB,
NOT to certify a letter ladder, and the T0 pathology is reduced rather than cured.
CANNOT DETERMINE: per-archetype calibration (Mandate 3d) — bucket n falls below
the reporting floor once split by sport AND archetype on 442 rows.
Queries committed at scripts/pwin-calibration-holdout.sql.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
STOPPED at the T0 gate as instructed. Nothing fixed, no recalibration applied,
no grade touched. T1-T4 deliberately not run.
T0 FIRES ON BOTH PRE-REGISTERED CONDITIONS.
Condition 1 (mean |predicted-actual| > 0.05): MLB ~0.094, WNBA ~0.139.
Condition 2 (monotonic slope): over-confidence GROWS with the prediction —
MLB +0.034 -> +0.043 -> +0.084 -> +0.190 -> +0.189; WNBA +0.044 -> +0.109 ->
+0.349. Worst cases: MLB predicted 0.842 actual 0.652 (n=23), predicted 0.917
actual 0.727 (n=11); WNBA predicted 0.730 actual 0.381 (n=21).
WHY THIS EXPLAINS THE INVERSION, mechanically: p_win is over-stated and the
overstatement SCALES with p_win, so p_win - fair_prob_lock is largest exactly
where p_win is most inflated. Those props hit less than claimed, so the edge
measure correlates negatively. The market was never the problem —
fair_prob_lock is not a bent ruler, the thing subtracted from it is. It also
explains why p_win ALONE still carries signal (+0.23 MLB): rank survives
miscalibration, differences do not.
This independently reconfirms the 2026-07-26 calibration finding (+0.02 at p<.5
-> +0.19 at p>=.8) on a newer, larger population, so it is structural rather
than sampling noise.
PART 0: P0a — only the GRADED side's fair prob is stored (fair_prob_lock;
no opposite-side field), so T1's two-side-sum check cannot run and must use the
stated no-vig recompute fallback. P0b — projection_locked_at exists as a
timestamptz so T2 is potentially runnable, but distinctness from lock time was
NOT verified because T0 gated it.
Two cautions recorded before Part 2 runs: the top MLB buckets where the error is
worst hold n=23 and n=11, so a flexible per-bucket correction would fit noise —
isotonic with pooling or single-parameter Platt is safer; and calibration fixes
magnitudes, so if the market is genuinely better the repaired edge may still land
at ~0, which would be the honest ceiling and gets reported rather than graded
around.
Query committed at scripts/grade-calibration-t0.sql.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
STOPPED at the Part 1 gate. Nothing rebuilt, no grade changed, no cutover.
THE FINDING: grading on p_win vs fair_prob does not work. All three candidate
edge formulations correlate NEGATIVELY with outcomes, on both sports, overall,
and in both time splits (n=432 decided rows carrying p_win AND fair_prob_lock):
ALL n=432 champ -0.0016 p_win ALONE +0.1221 additive -0.0615 ratio -0.1161 logodds -0.0438
MLB n=240 champ +0.0984 p_win ALONE +0.2278 additive -0.0336 ratio -0.1350 logodds -0.0124
WNBA n=192 champ -0.1143 p_win ALONE -0.0842 additive -0.1326 ratio -0.1281 logodds -0.1243
Subtracting the market's lock-time fair probability destroys and inverts the
signal. The plain reading: props where the model most disagrees with the market
are LESS likely to hit — the market is better than the model, so "edge vs market"
is anti-predictive here, while the raw probability retains some skill alone.
WHAT DOES CARRY SIGNAL: p_win alone, MLB only, and it is modest. Time-forward
split — TRAIN (07-21..07-26, n=120) r=0.2770; HOLDOUT (07-26..07-30, n=120)
r=0.1647, with the additive edge negative in BOTH halves. So p_win survives
forward validation directionally but the holdout is NOT significant (t~1.81,
p~0.07). Suggestive, not proven.
WNBA MUST ABSTAIN: every measure negative including p_win itself (-0.084). Forcing
one threshold across both sports would make a coin-flip sport look sharp, which the
order forbids.
LOOKAHEAD GUARD SATISFIED: fair_prob_lock is the lock-time field, populated on 432
decided rows, range 0.145-0.713. closing_prob (415 rows) is the CLOSE and was NOT
used in any correlation — using it would have manufactured a correlation.
SAMPLE REALITY: 1103 decided rows but only 432 carry both instrument fields, so a
per-sport train/holdout split leaves ~120 per half — enough to show direction, not
to certify a letter ladder.
I did not tune toward a win: three pre-registered candidates were tested and all
three failed; picking a fourth because the first three lost is the overfitting the
order guards against. Recommended instead: grade MLB on p_win alone with WNBA
abstaining and label it modest/accruing (A-RATED hold stays); or wait ~6 weeks for
n~500 MLB; or investigate WHY the market-relative edge inverts, which is the more
valuable question.
Both queries committed at scripts/grade-correlation-proof.sql so no number here
has to be taken on trust. Working settlement untouched; dead resolve endpoint not
wired.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Nothing built. No poller wired, no capture change, no env flipped.
PART 1 — GRADES ALREADY AUTO-SETTLE. snapshotScheduler resolves settleAllOutcomes
(:64) and settleAllLedgers (:67) and runs them FIRST at every snapshot slot before
grading (its own comment at :393, Session 61). The record is self-populating: 937
settled rows, growing daily (07-24 through 07-30: 20, 25, 44, 26, 98, 62, 91), and
/api/accuracy reads it live at 937 @ 58% (MLB 526 @62%, WNBA 411 @54%).
/api/grading/resolve is a separate unreferenced legacy path, not the settlement
path. Wiring an ESPN poller to it would create a SECOND settlement path racing the
working one and double-count an append-only ledger — so nothing was built.
The DNP/VOID requirement is already satisfied: outcome carries void and
unrecoverable as terminal states, and getModelAggregate excludes both from the
record denominator, so a DNP is never counted as a loss (105 void rows exist).
Idempotency is enforced too — settleLedger guards on .is('outcome', null) and
outcomeService dedupes on nameKey|stat|line|side|date.
THE REAL GAP is smaller and different: settlement covers MLB + WNBA only. NBA and
soccer grade but never settle because no free settled-result feed is wired. That
is a per-sport feed problem, not a missing poller.
PART 2 — clvCaptureReliable() is ONE LINE:
return process.env.CLV_CAPTURE_RELIABLE === '1';
It measures nothing. It fails because the operator has not set the flag, not
because the capture is unreliable. So there is no capture code to repair for the
guard to pass — flipping one env var publishes beat_close_pct immediately, which
makes this a judgement call and precisely the "make a number appear" move the
honesty guard forbids.
The guard itself works: beat_close_pct and clv_distribution publish only when the
flag AND settled>=20 AND clv_sample>0; with it off /record shows NOT PUBLISHED YET
and the computable 34/937 = 3.6% is never the publishing path (clvPanel returns
null and a test forbids the fallback).
CANNOT DETERMINE (Supabase MCP upstream-auth outage): the close-vs-locked
distribution, which is the direct test for the old silent-overwrite bug. The exact
query is in the report. A decision rule is stated BEFORE seeing the number so it
cannot be fitted to it: set the flag only if close_moved is a clear majority of
rows carrying a close AND coverage of settled rows is high enough that the
percentage describes the record rather than the captured subset. If either fails,
leave it off — a CLV near zero because close==locked is the fabrication to avoid
and it would look like success.
PART 3 — full outstanding board included in the report, covering model work
(A-flood grade fix on p_win vs fair_prob, the collapsed-output re-adjudication
list, calibration/time-series with no honest source, price-triplet MODEL leg),
surfaces (D1 mount, share cards, notifications, Offseason, /system, S3 media,
45 unwired glyphs, /record has no nav link) and infra (NBA/soccer never settle,
three credentials still flagged for rotation, migration drift 023-029).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Stripe wired to the Phase-A mechanism. Live prices verified READ-ONLY; no
Stripe object was created and no payment was run.
B1 PRICE KEY -> ID + BOOT ASSERTION (src/config/stripePrices.js). claim_founder_slot
returns a price KEY; this module is the only place a key becomes a Stripe id, and
it reads env (legacy STRIPE_PRICE_ANALYST/DESK accepted as fallbacks so an existing
deploy keeps working). assertPricesConfigured() is wired into server.js and FAILS
BOOT when any of the four is unset — verified by deleting one: it throws
"BOOT FAILED - unset Stripe price env for: desk_founder". A blank price can no
longer sell at the wrong rate or 503 a customer at checkout.
B2 CHECKOUT CLAIMS BEFORE CREATING THE SESSION. resolveCheckoutPrice previously
called founderSeatsAvailable() — a COUNT read, which WAS the race (two checkouts
at seat 99 both read 99, both got founder). It now calls claim_founder_slot and
uses the returned key. The promo-code bypass is retired: founderCode no longer
influences price or metadata, and getPriceId THROWS if handed a code rather than
silently granting a founder rate. metadata.is_founder is renamed is_founder_audit
and the webhook no longer reads it — caller-supplied metadata must never decide
who pays the lifetime founder price.
TRANSIENT-FAILURE POLICY (a real design call, not a default): if the claim RPC
errors we now fail RETRYABLY (503 claim_failed) instead of silently selling at
standing. Both silent options are irreversible — standing permanently overcharges
someone who was entitled to founder, and granting founder without a slot pushes
past the 100 cap at permanent prices. A full cap is NOT an error and still returns
standing normally, per "never error to the customer": a full cap is a real answer,
a DB blip is not.
B3 WEBHOOK FINALIZES THROUGH THE SINGLE WRITER. checkout.session.completed calls
finalize_founder_slot, which flips user_profiles.founder_pricing (canonical) and
mirrors users.founder_status in the SAME txn, so they cannot drift again (they
already had, 1 vs 0). Verify-after-write re-reads the profile and logs the end
state. If finalize errors, the tier is still set so a PAID customer is never left
unentitled, but no founder flag is guessed.
B4 SIGNATURE VERIFICATION was already present (constructEvent with
STRIPE_WEBHOOK_SECRET + express.raw). The live endpoint exists and is enabled:
https://api.vyndr.app/api/stripe/webhook subscribing checkout.session.completed,
customer.subscription.created/updated/deleted, invoice.payment_succeeded/failed.
VERIFICATION: V1 boot assertion proven by simulation. V2 all four prices retrieved
live and confirmed active with correct amounts and monthly recurrence (14.99 /
24.99 / 44.99 / 59.99) — read-only, nothing created. V3 no code path grants founder
except the claim (greps clean; the legacy helper now throws). V4 the handler reads
customer/subscription/metadata.user_id and calls finalize with signature
verification in place. V5 reset to a pristine 100 free / 0 claimed baseline with
both founder flags at 0.
Secrets live only in .env (0600, gitignored, untracked). A pre-commit scan
confirmed NO tracked file contains the key material.
Floor: 320 suites / 3984 passed, 3 skipped (superseded founder-code tests), web
build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
DB only. No Stripe call, no checkout/webhook rewire (Phase B). Migrations 035,
036, 037 applied to prod and tracked; repo files added.
035 SCHEMA TRUTH — user_profiles gains stripe_customer_id and
stripe_subscription_id (G3 proved the webhook stores neither today, yet
finalize and grandfather reconciliation both key off the subscription id), plus
a partial unique index so a subscription id resolves to exactly one profile.
036 THE MECHANISM — founder_slots is a real TABLE replacing the decorative view.
The claim is a single UPDATE whose target row is chosen FOR UPDATE SKIP LOCKED;
no count is read in the decision path. UNIQUE(slot_number) plus a PARTIAL
UNIQUE(user_id) WHERE status <> 'free' (one live slot per user). Seeded 100 free.
Q1 global pool: the slot travels with the user, so analyst->desk keeps founder
with no second claim. Q2: release_expired_slots handles TTL abandonment ONLY —
cancelled slots retire, so the counter only rises. A6 redirects
founder_pricing_seats to count claimed slots, capped 100.
PRICE IDS ARE NOT IN SQL. claim_founder_slot returns a price KEY
(analyst_founder / analyst_standing / desk_founder / desk_standing) and the Node
layer maps it to STRIPE_PRICE_* env with a boot assertion — adopted over
hardcoding so a typo fails at boot instead of becoming a permanent mis-charge.
A7 FLAG COLLAPSE — finalize_founder_slot is now the SINGLE writer of both
founder flags in ONE transaction: user_profiles.founder_pricing is canonical and
users.founder_status mirrors it. founder_status is NOT dropped (G5 proved it
live: written at stripeService:163, served at routes/stripe:95, loaded in
middleware/auth:24 PROFILE_COLUMNS). Only the independent write is retired —
the two flags had already drifted in prod (1 vs 0).
A9 RACE TEST, run in Supabase before any Stripe:
- pool squeezed to ONE free slot; three distinct users claimed concurrently
-> EXACTLY ONE is_founder=true on slot 100, two returned analyst_standing,
zero double-allocation.
- idempotency: the winner claiming again returned the SAME slot 100 and still
held exactly 1 live slot (two tabs cannot take two seats).
- constraint layer proven directly: a raw UPDATE granting that user a SECOND
live slot was REJECTED by the partial unique index, and verify-after-write
confirmed state unchanged (1 live slot, target row untouched).
HONEST LIMIT: the three claims contend within one transaction via LATERAL, so
this proves the claim logic, the SKIP LOCKED path and the constraint that makes
parallel safe — but it is not N genuinely parallel backend sessions. True
multi-session concurrency is not drivable through this SQL interface and should
be exercised once in Phase B against the test key.
037 NEXAPAY DROP — own migration, evidence-led (G4: zero code refs, column
empty). VYNDR is Stripe-only.
CLEAN BASELINE (Q3) verified after the test: 100 free slots, 0 non-free, counter
0/100, and BOTH founder flags cleared to 0 across user_profiles and users — the
inconsistent test record is no longer enshrined as a founder.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
No build, no migration, no Stripe object touched. Awaiting Kev on Q1-Q3.
TWO EXPECTATIONS IN THE ORDER ARE WRONG:
1. G5 — users.founder_status is LIVE, not dead. Written by the webhook
(stripeService.js:163), read and served by routes/stripe.js:95 as is_founder,
and present in middleware/auth.js:24 PROFILE_COLUMNS so it loads on EVERY
authenticated request. The guardrail says don't write it unless G5 proves it
live — G5 proves it live, so A5 must NOT drop it.
2. THE TWO FOUNDER FLAGS ALREADY DISAGREE IN PROD: user_profiles.founder_pricing
is true on 1 of 3 profiles while users.founder_status is true on 0 of 3. The
webhook writes both from the same isFounder, so this is a dual-write that has
already drifted. The build must pick one canonical flag and derive or retire
the other; two independently-writable founder flags is how a founder loses
their rate on one code path.
GREPS: G1 founder_pricing has exactly one writer (the webhook mirror) and four
readers (partners MRR attribution, the profile API, the profile badge). G2 the
promo-code bypass is the ONLY founder gate today — getPriceId(tier, founderCode)
against VALID_FOUNDER_CODES, stamped into metadata.is_founder, which the webhook
then trusts, so a code alone mints a founder at any seat number. G3 the webhook
DOES set tier + subscription_status=active + founder_pricing (closing an earlier
CANNOT DETERMINE: a paid sub does flip the Build-1 gate) but stores NO
stripe_subscription_id, confirming A1. G4 nexapay has ZERO code references and
the column is empty, so A5's drop is evidence-supported as its own migration.
G6 price selection is getPriceId -> line_items.
DB VERIFIED: user_profiles has nexapay_customer_id and NO stripe_customer_id /
stripe_subscription_id (A1 needed); users already carries stripe_customer_id;
founder_pricing_seats is a VIEW; 3 profiles, 1 flagged founder.
CANNOT DETERMINE: the four Stripe price IDs — no STRIPE_SECRET_KEY or
STRIPE_PRICE_* in this environment, so I could not independently re-verify that
the IDs in the order are what prod will charge. Since A3 would hardcode them, a
typo becomes a permanent mis-charge; recommend reading them from env (already the
pattern) with a boot assertion that all four resolve.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Nothing built. No Stripe object created or changed, no price logic touched.
STOPPED because there is no STRIPE_SECRET_KEY in this environment (.env holds
only ODDS/SUPABASE/INTERNAL keys). The order's standing floor requires
founder/standing/grandfather/race all verified server-side; none of that is
verifiable here, the standing price objects cannot be created, and the
concurrent-checkout race cannot be exercised. On a payment path the failure modes
are permanent and customer-facing — a race bug mis-prices a subscriber forever,
a grandfather bug overcharges one every month — so it must not ship unverified.
VERIFIED ANYWAY:
- Stripe IS live and FOUNDER price objects DO exist. /api/founders/count returns
{available:true, claimed:0, total:100}, and routes/founders.js returns
{available:false} whenever countFounderSeats() is null, which it is when
!STRIPE_SECRET_KEY || founderPrices.length === 0. So available:true proves the
secret key and at least one founder price ID are configured in prod, and
claimed:0 is a real count rather than a fallback.
- THE COUNTER IS NOT A GATE. It is a cached (300s) READ, not a claim; founder
pricing is gated by CODE + EXPIRY, not by the count, so anyone holding
FOUNDER2026 gets the founder rate at any seat number and the cap is decorative.
Two simultaneous checkouts at slot 99 would both read 99 and both get founder —
there is no lock or unique constraint anywhere in the path.
- The gate reads users.tier via config/tiers.js reasoning_visible, so a
successful subscription must set users.tier for Build 1's gate to open.
CANNOT DETERMINE: whether the STANDING price objects exist (env unreadable, and
getPriceId falls back SILENTLY to a PRICE_UNCONFIGURED sentinel, so a missing
standing object would not surface until the first post-cap checkout 400s in front
of a paying customer); whether the webhook writes users.tier on
checkout.session.completed.
DESIGN IS SETTLED for when it unblocks: a founder_slots table with a unique
constraint on (tier, slot_number) claimed before the Stripe call — the unique
index, not a count read, is what makes the race impossible; price selection from
the claim rather than a code, with the code+expiry bypass retired; grandfathering
by simply never calling Stripe price-migration on a founder sub;
founder-follows-upgrade by claiming on the target tier and releasing the slot on
cancellation; honest display that shows no number when the count is unavailable
(the existing route already sets that precedent).
PREREQUISITES, all needing Kev and none of them code: confirm/create the two
standing price objects and set STRIPE_PRICE_ANALYST / STRIPE_PRICE_DESK; confirm
the webhook sets users.tier; provide a Stripe test-mode key so the race,
grandfather and end-to-end unlock can be exercised rather than asserted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Presentation over existing endpoints. src/ untouched (git diff empty): no grade,
model or ledger change. Pricing/migration are Builds 2/3.
TIER-RECORD-FORWARD. /record reads the canonical public aggregates (/api/accuracy
+ /api/ledger/model) and prints them as-is. B 60% n512 and C 57% n413 ship with C
honestly BELOW B; A (n2), D (n5) and F (n5) render HOLLOW with their real sample
instead of a rate. Sport slicing (all/mlb/wnba) is client-side because the
endpoints ignore ?sport= — mlb 526 @62%, wnba 411 @54% come from the sports map.
THE LOAD-BEARING RULE, enforced in lib/proofRecord.js and locked by tests: where
the source withholds a percentage it stays null. A is 1/2 and therefore 50% is
derivable — a test asserts we do NOT derive it, because the API withheld it on
purpose (n < 20).
CLV IS AN HONEST ABSENCE, NOT A NUMBER. beat_close_pct is null because
clvCaptureReliable() has not passed. The panel says "NOT PUBLISHED YET" and
explains that any percentage printed today would be measuring our collection gaps
as much as our edge; it surfaces the accruing sample (937) but no rate. Tests
assert the panel never falls back to clv_beat/clv_sample (34/937 = 3.6%) and that
the serialized panel contains no "3.6" — that number is computable and would be
wrong, which is the exact fabrication this surface exists to refuse. The panel is
built to receive a real number later without a redesign.
HELD, and named on the page rather than faked: calibration and accuracy-over-time
are absent because there is no honest source (no claimed-vs-actual endpoint;
window_days fixed at 30 with no series). The page says so, and says it is not
because they are unflattering.
A page-level test asserts no hard-coded percentage exists in the markup, so no
figure can drift from the aggregate, and that the page never touches /api/snapshot
or itemized rows — the Build-1 gate holds and the exploit stays dead.
Floor: 320 suites / 3986 tests green (16 new), web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Nothing built. Docs only. The premise was that this is cheap assembly over
existing aggregates; verified, it is not.
0.1 FILTERABILITY — the endpoints are NOT filterable. Probed live:
/api/ledger/accuracy?sport=mlb -> total 937
?sport=wnba -> total 937
?window=7 -> total 937
identical payloads; the params are ignored (the req.query reads at
routes/ledger.js:61-63 belong to a different route than /accuracy at :68).
Sport and tier CAN be sliced client-side from /api/accuracy's sports map and
/api/ledger/model's by_tier. TIME WINDOW CANNOT — window_days is fixed at 30
inside getModelAggregate with no param and no stored series, so
"accuracy over time" has no data source.
0.3 CLV CANNOT LEAD WITH A NUMBER. /api/ledger/model exposes the aggregate, and
live it returns beat_close_pct = null and clv_distribution = null despite
clv_sample 937. They are null BY DESIGN: ledgerService publishes them only
when clvCaptureReliable() passes, and it does not — the capture is still the
starved instrument the 07-28 repair improved but did not finish. The trap to
avoid is exact: clv_beat/clv_sample = 34/937 = 3.6% is computable and would
be WRONG, because the value is null due to instrument distrust, not a missing
division. Publishing it would be the marketing fabrication this order most
forbids. CLV can only lead with an honest absence.
0.2 The honest-record laws are ALREADY enforced at source: buckets return
A pct:null (n=2), B 60% (512), C 57% (413), D pct:null, F pct:null — thin
tiers already refuse to round. C genuinely sits below B, which is the
unflattering truth and must be shown as-is.
0.4 CALIBRATION CURVE has no data source — clv_distribution is null and there is
no claimed-vs-actual endpoint; the 07-26 calibration work was a one-off
read-only measurement, never wired to a served surface.
BUILDABLE NOW: tier hit-rates by sport with existing hollows preserved,
client-side sport/tier filtering, the capped 3-call sample, and honest state copy
including a CLV not-yet-publishable panel that names the reliability guard.
NEEDS ITS OWN ORDER FIRST: CLV as a leading number (blocked on capture
reliability, not presentation), accuracy over time (needs a param or daily
series), the calibration curve (needs a claimed-vs-actual endpoint).
RECOMMENDS shipping tier-record-forward with an honest CLV building panel rather
than CLV-forward — CLV-forward with a null cannot lead, and with 3.6% would be a
lie. That preserves the premise's strongest claim (a real thin honest record
out-credibilizes a fake fat one) without inventing a number.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Serving/gating change only. src/services/ untouched: no grade, model or
settlement-logic change. Pricing = Build 2, migration = Build 3.
WHY THE PRIOR GATE WAS WRONG: freeing grades at resolution made the free tier a
ONE-DAY-DELAYED FEED OF THE WHOLE PRODUCT — settlement is nightly, so a bettor
watching one cycle behind got the entire method free. There is now NO
per-grade resolution flip: an itemized grade, tonight's or last week's, is
Analyst+.
FREE now gets, none of it itemizing the nightly slate:
1. the full data aggregator (unchanged — schedule, per-book lines, stats,
streaks, hubs)
2. the AGGREGATE track record, which ALREADY EXISTS and is public:
/api/accuracy (sample 937, byGrade tiers, per-sport mlb+wnba, min_sample 20)
and /api/ledger/accuracy (per-grade buckets). The honest-record laws are
already honored there — A/D/F return pct:null under the n>=20 threshold
rather than a fake percentage.
3. a CAPPED, day-rotated sample of resolved calls for texture: cap 3, stable
within a day, rotates across days, and only RESOLVED rows are eligible so a
live read can never be sampled. The cap is what kills the exploit — three
rotating past calls cannot reconstruct a nightly slate, whereas the full
settled list is the feed one cycle late.
4. the locked shell of tonight's reads: they exist, and their shape.
EVERY itemized grade for an unentitled tier now loses grade, confidence,
confidence_basis, reasoning, kill_conditions_triggered, projection, edge_pct,
matchup_grade, form, alt_lines and kelly, and is stamped locked. Free-side DATA
survives so the board still reads as real: player, market, line, book_odds,
fair_odds (the de-vigged fair number is the free hook and is never the paywall),
season/last10 stats, archetype — and `outcome`, because a RESULT is a fact
rather than a judgment.
The tease stays aggregate-only (live_locked {count, tiers}) computed from the
ungated rows and never joined back to one, and no gated row carries a grade, so
nobody can work out which prop is the A.
Floor: 319 suites / 3970 tests green, web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Serving/gating change only. src/services/ untouched (git diff empty): no grade,
model or settlement-logic change. Pricing and migration are Builds 2 and 3.
Push scoring untouched.
THE RULE: a grade is PAID while its outcome is unknown and becomes FREE the
moment it resolves.
Resolution is read ONLY from a written outcome — never from time, game status or
gradedAt. A game can be final long before the settle pass runs, so treating
"probably over" as settled is exactly how a live edge would leak; a test asserts
an hours-old gradedAt with no outcome is still LIVE. void and unrecoverable ARE
resolutions (terminal results, no live edge left). isResolved FAILS CLOSED:
null outcome, {} with no result, and empty-string result all read as LIVE, so a
settlement failure withholds content rather than exposing it — the same
direction resolveTierFromRequest fails.
FREE/ANON: settled grades pass through IN FULL, reasoning and kill conditions
included — settled reads are the proof product and cost nothing once the outcome
is known. That also converts the previously-unenforced board reasoning leak into
a deliberate rule rather than an oversight.
LIVE grades for unentitled tiers are reduced to a shell: every piece of model
JUDGMENT is dropped (grade, confidence, confidence_basis, reasoning,
kill_conditions_triggered, projection, edge_pct, matchup_grade, form, alt_lines,
kelly) and `locked: true` is stamped so the card renders the unlock prompt. The
free-side DATA stays so the tease is real rather than empty: player, market,
line, book_odds, fair_odds, season/last10 stats, archetype, gradedAt, history.
fair_odds deliberately survives — the de-vigged fair number is the free hook and
is never the paywall. A test asserts the serialized free row carries no trace of
the withheld judgment.
THE TEASE IS AGGREGATE ONLY: live_locked = {count, tiers} computed from the
ungated rows and never joined back to one, and no gated row carries a grade — so
a free viewer learns that N reads exist and their tier shape without being able
to work out WHICH prop is the A.
Gate order in the route: stripModelPrice (S67) first, then gateLiveGrades.
Entitled tiers get the array back by reference — zero cost, zero change.
Floor: 319 suites / 3971 tests green (10 new), web build exit 0.
One test note: the route-level supertest case was removed deliberately — it
needs a live Redis and hangs on ioredis' reconnect timer in a single-suite local
run (known behaviour, CLAUDE.md). The gate contract is fully covered by pure
tests; the wire is verified against prod anonymously in the fingerprint.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Report-first. Nothing built; no tier, price, gate or Stripe object changed.
REVIEW ZERO findings that shape the design:
0.2 The ladder is HALF-EXPRESSIBLE already — PRICE_MAP separates founder from
standing objects, so lifetime grandfathering is native (a sub created against
a founder price stays on it). BUT founder access is gated by CODE + EXPIRY
(FOUNDER2026/VYNDR/BETONBLK/EARLYBIRD, expiry 2026-12-31), NOT by seat count:
anyone with a code gets founder pricing at any seat number. A real
Stripe-derived counter exists (/api/founders/count, live 0 of 100) but only
DISPLAYS — and it is cached 300s, so it cannot enforce "slot 100 and 101
differ permanently". Making the counter the gate, transactionally and
uncached at checkout-session creation, is a real build.
0.3 The paid->free flip point already exists ON THE SERVED PAYLOAD: settlement
writes ledger_entries.outcome + settled_at, and /api/snapshot already merges
per-grade results — live WNBA returns 25 grades, 5 carrying
outcome {result:'hit', actual:1}. So the gate discriminator (outcome != null)
is present on the exact object to be gated; no new pipeline needed.
0.4 THE MIGRATION IS NOT WHAT THE ORDER ASSUMES: the users table holds 3 users,
all free, created Jun 12-19, and ZERO paid. There is no warm mass base — the
"founder launch to existing users" is a courtesy note to 3 people, and the
launch's real audience is people who have not signed up yet.
DESIGN: free = full data aggregator + the COMPLETE settled record (letter,
reasoning, edge, outcome — browsable and filterable), which is the proof hook.
Analyst = tonight's live grades + reasoning + edge, unlimited. Desk = + alt
ladder, Kelly, portfolio, engine2. Reasoning/grade/edge are ONE paid unit while
live and become free together at resolution — which also converts today's
unenforced board-reasoning leak into a deliberate rule.
GATE: outcome == null => live => Analyst+; outcome != null => settled => free.
Filter whole grades server-side (not field-strips) so a live grade cannot leak
partially; never infer resolution from time or game status, only from a written
outcome; fail closed to LIVE so a settle failure withholds rather than exposes;
void/unrecoverable are terminal and therefore free.
BUILD ORDER: (1) the settled/live gate, (2) the free settled-record surface —
noted as arguably shipping WITH (1), since gating live grades without it leaves
free users no graded content at all, (3) Stripe ladder + transactional counter +
grandfather rule + retire the code gate, (4) the founder note to the 3,
(5) pricing visuals (already designed in the package).
CANNOT DETERMINE: whether the four Stripe price objects exist in the dashboard
(env not readable here) — flagged as a prerequisite for build 3.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Read-only. Nothing changed.
FREE TIER, EXACTLY:
- Board /api/snapshot: NO count limit. The only gate is
stripModelPrice(grades, tier) at routes/snapshot.js:106-107 — no slice, no
volume branch. Live anonymous right now: MLB 5, WNBA 25 = the full board.
The "3 scans/day" cap rations the SCAN path only.
- Grade letter: fully visible on every tier (grade_visible: true). Anon also
receives confidence, edge_pct and VYNDR's own projection.
- Edge fields: correctly stripped. p_win/ev_pct/model_odds/value/takeable are
ALL absent from the anonymous payload, with model_price_locked stamped so the
card shows a lock teaser rather than an absent leg. This half works as designed.
THE HEADLINE — the two paths disagree on reasoning:
- Scan REDACTS it: tierGating.js lockReasoning + lockKillConditions +
tier_gated + upgrade hint, driven by free.reasoning_visible = false.
- Board SERVES IT IN FULL: snapshotGating MODEL_FIELDS is
[model_odds, p_win, ev_pct, value, takeable] — reasoning is not in the list.
Verified live anonymously: full reasoning.summary plus a kill condition WITH
its reason.
Intent: config/tiers.js declares free: { reasoning_visible: false } with the
comment "blurred — frontend renders tier-locked". One of the two paths does not
enforce the product's own declared line, so the evidence reads as oversight
rather than funnel — a funnel would be declared in config, not contradicted by
it. Flagged with the counterweight: board reasoning is good marketing and the
data layer is already free, so closing it is a monetization tightening (Kev's
call), not a fabrication fix.
FREE DATA IS A REAL AGGREGATOR, not just a limited graded view: schedule,
per-book lines, player stats, streaks, hot lists, team hubs, public record — all
public and uncapped (probed live).
PAID (config/tiers.js, checkout.js:4): analyst $14.99 / desk $44.99. Analyst is
unlimited reads; Desk differentiates on capability (alt ladder, Kelly, portfolio,
engine2). africa tier is defined but activation is blocked on a DB CHECK
constraint. api_access is false on every tier. book_odds/fair_odds deliberately
pass through on all tiers — the de-vigged fair number is the hook and is never
the paywall; only model_odds gates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Nothing changed: no mount, no row edit, no data threading. Docs only.
1. THE RATIONALE DOES NOT REACH THE ROW. StripProp carries stat/line/side/grade/
gradedAt/delta/awaiting/outcome/movement/revisedFrom/book/bestBook/dead/
history — no reasoning, no kill_conditions_triggered — and
buildPlayerStripsFromProps never threads them. Mounting the hover needs a new
field on the strip contract threaded through the slate adapter: additive, but
a data-path change rather than a mount.
2. THE 0.3 PREMISE INVERTS — THE RATIONALE IS ALREADY PUBLIC. Verified live and
anonymously against prod: /api/snapshot/wnba returns reasoning.summary with no
locked flag plus kill_conditions_triggered. stripModelPrice removes
model_odds/p_win/ev_pct/value/takeable but NOT reasoning. So the full model
rationale already ships to every anonymous browser on the main board, while
the same content IS tier-gated on the scan path (tierGating.js). Mounting the
hover would leak nothing new, but would surface content that is currently
shipped-but-unrendered, and the product gates it in one place while serving it
openly in another. That is a monetization/consistency decision, so it is
reported with three options rather than resolved unilaterally.
3. ROW-GRAMMAR IS LAW AND LOCKS StatStrip's SOURCE ORDER. rowGrammar.test.js
asserts element order via src.indexOf on the component source; adding a
rationale affordance or a team chip moves those offsets, so specs/ROW-GRAMMAR.md
and the test must be amended in the same commit. That makes this spec-amending
work needing its own slot decisions, not an additive mount.
Safely mountable with no blockers: reveal.js (wraps the row list, no StatStrip
internals, no new data, no grammar slot). teamChips needs a grammar slot;
rowRationale needs the data threading AND the gating decision AND a slot.
Recommends splitting D1-close into: mount reveal now; a ROW-GRAMMAR amendment
order for the chip + rationale slots; then the rationale mount once the gating
decision is made.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Additive frontend. Backend untouched (git diff src/ = empty): no grade, model,
classifier or ledger change. Scope held to the row anatomy these three items
need — no System-artboard-wide rebuild. Push scoring untouched.
REVIEW ZERO — the two checks that decided whether these could be honest:
0.2 RATIONALE SOURCE — VERIFIED REAL. Live snapshot grades carry `reasoning`
and `kill_conditions_triggered`. The summary is built by analyzeViaEngine1
from the actual feature vector (l5/l20 averages, gap to the line, home/away,
opponent defensive rank, rest days) and kills carry real codes + reasons.
So the hover shows genuine grade truth, not a placeholder.
0.3 TEAM COLOURS — PARTIAL, and deliberately left partial. The System artboard
defines a colour pair for only 10 teams (BOS CHC CHI DEN LAD MIL MIN NYY PIT
SD), lifted verbatim; lib/teams.js holds ~80. The other ~70 are NOT invented
— a wrong team colour is a recognition error the user reads as fact. Unknown
teams get the honest-neutral chip (muted border, no colour claim), never a
guess and never a blank gap. Coverage is reported by coverage(), not hidden.
SHIPPED:
- web/src/lib/rowRationale.js — rationaleFor() returns real summary + kills, or
NULL. No generic fallback: an empty hover is honest, a manufactured "why" is a
fabricated model explanation. A locked/tier-gated reasoning is treated as
ABSENT rather than paraphrased or leaked, and a kill condition with no reason
explains nothing so it is dropped.
- web/src/lib/reveal.js — IntersectionObserver reveal that fires ONCE then
unobserves ("react to truth, then rest"), reuses D1-A's bootDelayMs for the
60ms stagger so there is ONE source of truth for the timing, and reveals
IMMEDIATELY when IntersectionObserver is absent (SSR/test) so a missing API can
never hide real content. Reduced motion is handled by the existing CSS, so the
row is visible either way.
- web/src/lib/teamChips.js — Rev-3 geometry (10px, 135deg, before the abbr,
inside the row) plus the ranked opacity ramp 1/.86/.64/.48 so chips dim with
their row. Swap-ready for licensed logos at the same size.
Floor: 318 suites / 3961 tests green (15 new), web build exit 0.
The three modules are pure and unit-locked; mounting them into the live row
components is a follow-up, and the visual result belongs in the Chrome audit.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Additive frontend/visual. Backend untouched (git diff src/ = empty): no grade,
model, classifier or ledger change. The 41->74 registry expansion is HELD for
D1-B. Push scoring untouched.
REVIEW ZERO — classifier coverage bounded the glyph wiring. Three buckets, and
the computation was redone three times before it was right (the frontend keys
GLYPHS by ARCHETYPE NAME while the backend keys `glyph:` by SHAPE NAME, and most
registry keys are unquoted identifiers — the first two passes mis-parsed both):
(a) classifier-backed, already wired: 38
(b) classifier-backed, package SVG exists, NOT wired -> WIRED HERE: 6
striker, grappler, pressure, counter, grinder, finisher — all MMA/combat
archetypes in archetypeService.js that were rendering EMOJI fallbacks
('*', 'x', '>', '<>') where the package ships real 24-grid duotone marks.
(c) package SVG with no classifier -> HELD for D1-B: 39 (wiring them would
render nothing)
(d) classifier-backed but NO package SVG: 2 ('dual threat', 'paint boss') —
a DESIGN gap, not a build gap; flagged for D1-B.
GLYPHS map 38 -> 44 keys, deliberately far short of the package's 83.
BOUNDARY CHANNEL — the blue tokens already existed (--priced-out set) and were
applied on NoMarketState and the scan void box, but PriceTriplet's NO_MODEL
("line not priced") still rendered in neutral text, so the channel was applied
inconsistently. NO_MODEL now renders in the channel, completing "every boundary
state or none". Token-only (no hex fallback and no hex in prose — PriceTriplet's
own test forbids literal hex, and it caught both).
REACTION PRIMITIVES — new web/src/lib/reactions.js + globals.css keyframes at the
exact HANDOFF timings: flash .75s ease-out, boot stagger 60ms steps, reactions
gated 1.5s, WIRE hold 6s. nudge() REFUSES a no-op (null/absent direction -> no
flash) so the primitive cannot be attached to an idle loop — a flash without a
new datum is the UI lying about the feed. Reduced-motion honoured.
READ-FAB — aligned to the exact package geometry: 50px circle, translateY(-14px),
6px void ring (was 46px, marginTop -16, 3px ring).
CARD TOKEN — audit correction: #0E0E14 was already tokenised as --bg-1/--card;
the audit's "1 file" was counting the raw hex, not the token. No change needed.
Floor: 317 suites / 3946 tests green (16 new), web build exit 0.
NOT DONE THIS ORDER (reported, not silently dropped): row-hover rationale and
IntersectionObserver reveal (Phase 3 item 8) and team-gradient chips (Phase 4
item 10) are not implemented — they need the System artboard's row anatomy,
which is a larger port than the rest of D1-A.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Package specs/design-reference (Jul 22) audited against the CURRENT repo
(bf7c0a3, ~9 days later). No ~/vyndr_design exists; the in-repo copy is the
package. Every claim is a direct file/grep/count check, not the harness that
returned a silent false in Wave 3.
61 implementable items enumerated: BUILT-TO-SPEC 20, BUILT-BUT-DRIFTED 7,
PARTIAL 16, ABSENT 18 (+1 CANNOT DETERMINE: 19-screen mobile parity needs a
visual pass).
Largest single gap: the glyph library — 38 of 83 designed SVGs are wired (46%),
and the design implies 74 display archetypes against a 41-entry backend
registry, so the archetype system is roughly half the designed scope.
Drift found on surfaces built recently: the book comparison wired 07-29 renders
per-book lines but has NO crown, NO disagreement axis, NO SPLIT chip and NO
movement strip — a simpler version than the S2 design. The mobile tab bar has 5
tabs but not the designed READ-FAB. Calibration gating disagrees with the design
(our n>=20 vs designed N30).
Wave-2 reclassification: Newsletter DESIGN EXISTS (S5 The Report is fully
designed) — the earlier status pull was wrong to call it a design gap. Live
tracking and Slip reader remain genuinely design-missing.
Model linkages named: Price Triplet waits on the EV layer producing
p_win/ev_pct/model_odds; the S4 calibration curve waits on the n-threshold
decision plus accrued buckets, while the CLV chips can build on the repaired
instrument now.
Ordered build list in six dependency waves: self-contained first (glyphs,
primitives, boundary-channel blue), then scanner-nudge-gated, model-gated,
resolution-pipeline-gated (share-card masters cannot ship — the tail has no
generation step and no trigger), licensing-gated (book logos, push-to-book),
then the large surface builds.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
No grade, ledger or scoring change. Push scoring untouched.
REVIEW ZERO 0.3/0.4 — THE RESOLUTION TAIL DOES NOT FIRE. The resolver is
POST /api/grading/resolve (routes/grading.js:208), and its fanout at :356-371
covers webPush, telegram and discord — but:
- share-card generation: SPEC'D-NOT-BUILT. Not in the fanout at all (grep
shareCard in grading.js = 0). shareCards/renderer.js exists with ZERO
callers, so the component is built but no step would ever invoke it.
- push notifications: BUILT-NOT-FIRING. In the fanout but gated on
webPush.configured() (VAPID). push_subscriptions = 0 rows and
user_notifications = 0 rows — nothing ever subscribed or delivered.
- Telegram result posts: BUILT-NOT-FIRING (gated on BOT_TOKEN + CHANNEL_ID).
- Discord result posts: BUILT-NOT-FIRING (gated on webhookFor('results')).
- recap (all-Final trigger): SPEC'D-NOT-BUILT. No recap file exists in src/.
AND THE WHOLE TAIL IS UNREACHABLE: nothing calls /api/grading/resolve — there is
no ESPN poller in the repo. The live settlement path is the scheduler's
settleAllOutcomes + settleAllLedgers, which fans out to opsNotify only (ops
alerts), with no user-facing output. So even the built channels have no trigger.
Per the order's own rule, ShareCard, /notifications, result posts and recap are
therefore ALL SCOPED, none shipped — no dead shells over a silent pipeline.
BUILT — /compare. Semantics (0.2): a same-market head-to-head, two players with
every row a measure BOTH sides are scored on, aligned via alignRows so the
numbers are comparable — deliberately not two disconnected graded props. Reads
the live /api/stats/player/:name?sport= aggregate. Honest-absent three ways: an
unresolved side reads NO DATA while the other still renders; a measure only one
side has renders a dash, never 0; if neither resolves the page refuses to
compare. NO VERDICT — it shows measures and says the reader draws the call.
Two pre-existing tests (vyndrPhaseE, vyndrParityQA) asserted the in-development
placeholder; both superseded rather than deleted — they now assert the stronger
properties against the real page (live fetch, no sample players, NO VERDICT,
NO DATA, "not a zero").
Floor: 316 suites / 3930 tests green (10 new), web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Wiring + one copy pass. No grade, ledger, model or scoring change (diff empty
across intelligence/, ledgerService, outcomeService, gradeSlateService).
REVIEW ZERO — each surface proven with real data BEFORE wiring:
0.1 /intelligence vs /system are NOT duplicates. System.dc.html is a
multi-surface artboard (TERMINAL + INTELLIGENCE + WIRE sections), not the
design for a distinct /system route; its INTELLIGENCE section is already
realised as the live app/intelligence/page.tsx. No /system page exists and
none should be built as a second copy — the prod 404 is correct.
0.2 /intelligence renders live and gates SERVER-side, not by blur: the proxy
requires auth and limits by tier (desk 50 signals / non-desk 8), and
returns 401 to an anonymous caller (verified live). No leak.
0.3 /slip parses a real DraftKings slip end to end: 3/3 legs,
needs_review false, Aaron Judge total_bases over 1.5 @ -115. Honest limit
recorded: parsers are layout-rigid, an unsupported layout yields ZERO legs
rather than wrong ones (never-guess), so real-world OCR hit-rate across
layouts is CANNOT DETERMINE until user slips arrive.
0.4 /parlay direct route hits the real correlation builder on the same
ParlayContext the drawer uses.
0.5 /marketplace advertised four unbuilt things but made NO performance or
profit claim, and its capture was already real (/api/waitlist upserts to a
waitlist table). The gap was tense, not fabrication.
WIRED: Nav MORE gains Intelligence, Slip Reader and Marketplace; Parlay Lab
re-pointed from the drawer hash to /parlay (the drawer is unaffected —
ParlayPanel stays mounted with its floating badge).
GATING: /intelligence added to GATED_ROUTES because its feed 401s signed-out, so
an ungated link would land visitors on a permanently empty page. /parlay stays
OPEN deliberately — it is the free parlay funnel and gating it would be a
monetization regression.
/marketplace honesty pass: every item body now opens "Not built yet." /
"Not written yet." / "Not produced yet." with what is planned; the subhead states
it is not a purchase, not a pre-order and not a promise of a ship date; the
playbook item carries "No profit claim, no promised return". The capture stays
real — no fake button. Unit-locked.
Floor: 315 suites / 3920 tests green (12 new), web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Report-only. Nothing built, wired, tagged, or removed.
PART 1 — the recalibration boundary was NOT written. The build order asks to tag
grades pre/post edge-shading at "the true promotion timestamp"; there is no such
timestamp, and writing the marker would insert a fabricated model transition
into an append-only public record — the exact corruption the order exists to
prevent. Three independent production proofs:
A. model_snapshots.code_sha — every sha that ran the pipeline in the last 5
days is a documented commit from this session (f3bf300 currently live,
then f310608, 9b5235c, b8ee216, afb56b1, 3592aba, 914a057), all stamped
model_version engine1@2026-07-20. No promotion commit exists.
B. Daily A-family share 07-24..07-30: 0.0, 0.0, 0.0, 0.0, 1.0, 1.6, 0.0 —
flat at zero, no step change on any date. A 92.9% re-letter under shading
would have driven the board to ~79-93% A overnight.
C. efficiencyShading has zero production importers; HEAD d54eca0, clean tree.
Phase 2 delta: neither 92.9% nor 43.6% is attested in any measurement here, and
no re-lettering occurred at any scale, so partial-slate-vs-full-board cannot
explain a gap that does not exist. The only measured numbers, on all 1250 rows
(the full board): 97.4% would change, 79.8% up, 79.0% A-family, and 0 of 1250
rows actually shaded — the hypothetical re-letter would have come entirely from
an unapproved grading-basis switch, which is why the flip was refused.
Measurement integrity needs nothing new right now: model_version already
separates the S64 eras and takeable tags are complete (1245/1250). The boundary
becomes a hard prerequisite the day a promotion actually ships.
PART 2 — build triage for 14 incomplete surfaces, classified from code with each
one's real dependency and honest size, sequenced into five waves: pure wiring
(/intelligence, /slip, /parlay, /marketplace after a copy honesty pass);
design-only gaps (Live tracking, Slip reader, Newsletter artboards);
self-contained builds (/compare, ShareCard host, /notifications); model-gated
(price triplet MODEL leg and the calibration board both need p_win vs
fair_prob — building either on edge_pct would re-ship the retired 620% lie);
and sport-boundary/quota-gated (/soccer blocked on odds-api 0/500, /system,
Offseason).
Matrix corrections: Live tracking, Newsletter and Slip reader are built/live/
honest and marked incomplete only on the DESIGN column; /parlay is
drawer-reachable, not unreachable; /compare is already honest. No live surface
is showing fabricated data today.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Verified three ways that nothing was promoted and no re-lettering happened:
HEAD is the no-flip commit with a clean tree, efficiencyShading is imported by
zero production files, and live grades carry 0.0% A-family (MLB {B:1,C:4},
WNBA {B:15,C:10}). Neither 92.9% nor 43.6% is a figure measured here - the
challenger's real numbers were 97.4% would-change / 79.8% up / 79.0% A, with
0 of 1250 rows actually shaded.
Takeable tags landed (1245/1250 tagged, takeable_floor on all, one floor -160;
the 5 untagged have no locked price). The model-version boundary is absent and
correctly so - there was no recalibration to mark. Not a pre-audit gap.
Matrix re-derived from live prod probes, importer counts and nav-link counts:
15 of 26 fully done. Book comparison RESOLVED (BookComparisonPanel routed to
the grade card). Grade card honesty improved by the edge_pct retirement.
ShareCard / MobileEdgeBoard / DemoScan still dead code. Seven live routes
remain orphaned with zero nav links; /system and /offseason are 404. No
LIVE-but-not-HONEST surface found.
Design is NOT complete: Live tracking, Slip reader and Newsletter are shipped
with no design artboard - design is the gap, not build. Offseason is the
reverse (designed, never built).
Chrome audit manifest assembled: 11 items with per-item session state, four
requiring an entitled Desk session that only Kev can drive.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
THE PROMOTION WAS NOT PERFORMED. Champion grade path byte-identical (diff empty
across intelligence/, gradeSlateService, snapshotService). Projection, p_win and
the CLV instrument untouched.
REVIEW ZERO IS A GATE AND THREE OF FOUR PREREQUISITES FAIL:
0.1 scores are ESTIMATED priors from the founding spec, not measured. The
premise's cited values are not in the code either — the module holds
nba:points .80 and mlb:total_bases .55; there is no NBA 0.72 and no WNBA
score at all.
0.2 VERSION-BOUNDARY TAG DID NOT LAND — config/modelEras.js has zero shading
references. It was deliberately not applied twice (nothing had been
promoted) and reported both times. The order's own rule says STOP.
0.3 NO ROLLBACK FLAG EXISTS — zero occurrences of SHADING_ENABLED /
EDGE_SHADING / shadingEnabled anywhere in src/.
0.4 takeable tags DID land (migration 034, 1246/1254 rows). PASS.
AND THE APPROVED DELTA DOES NOT MATCH THE MEASURED ONE. Approved: 43.6% of
grades re-letter, efficient markets tighten and soft hold. Measured on all 1250
live rows: 97.4% change (1217), 79.8% move UP, 17.6% down, resulting in 79.0%
A-family (MLB 93.4%) against the champion's 0.2%. And rows_actually_shaded = 0
of 1250 — 96.5% of markets are unscored (f=1) and the one scored market present
is the anchor (f=1.0 by construction). The entire re-letter comes from switching
to edge-vs-fixed-bar grading, NOT from efficiency shading, which is inert on
this board. That is an unapproved grading-basis change riding along, which the
order's own "no new scaling changes riding along" guardrail forbids.
Flipping would re-letter 97.4% of an append-only public record, move 79.8% of
grades UP and mint A's on 79% of the board, on a letter whose measured
correlation with outcomes is r ~ 0.005 — the exact scenario the permanent
founder ruling forbids.
SHIPPED — ORDER B (independent of the promotion, and a live falsehood):
edge_pct display retired from GradeResultCard (confidence strip, EDGE stat cell
now honest-absent, alt-ladder rung) and SoccerGradeResult. DeskShowcase kept
(already honest). Computation and the board's signed-edge sort fallback SURVIVE
— deleting them would re-break the sort fixed on 2026-07-29; a test asserts all
three survive and the sort still orders agrees -> disagrees -> absent.
Fixed two build-breakers the retirement caused (orphaned edgeColor import,
orphaned edge_pct destructure; edge_pct stays on the props contract). Two
pre-existing tests superseded rather than deleted: they asserted the edge figure
is sign-coloured, and now assert the stronger property that no edge percentage
renders at all.
Floor: 314 suites / 3908 tests green (9 new), web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Challenger only. Champion grade byte-identical (verified by diff). Nothing
promoted, no live grade re-lettered, no ledger row deleted or re-settled.
BUILT src/services/challengers/efficiencyShading.js (measured-never-served):
adjusted_edge = raw_edge * f(efficiency); grade = band(adjusted_edge) against
ONE fixed bar (A+>=10, A>=5, B>=3, C>=1, D>=0, F<0) that never moves.
f(e) = E_SOFTEST/e bounded to (0,1] — soft markets intact (never amplified),
sharp shaded toward but not past zero, unscored -> f=1 and FLAGGED.
A fence test asserts no production grade path imports it.
Cross-market behaviour is unit-proven: the same raw 6% edge grades A in soft
mlb:total_bases and B in sharp nba:points.
MEASURED on 1250 live ledger rows — Phase 2.5's answer is NO, the flooding is
not gone: challenger 79.0% A and 80.9% A/B (MLB 93.4% A) vs champion 0.2% A.
TWO findings explain why, and they are the point of the order:
1. The shading is a NO-OP on the live board: rows_actually_shaded = 0 of 1250.
96.5% of rows are UNSCORED (f=1), and the one scored market present
(mlb:total_bases) is the anchor so its f is 1.0 by construction.
mlb:strikeouts and nba:points do not appear in the ledger at all (our
basketball is wnba, not nba). Challenger vs baseline: 0 rows changed.
2. Placement was never the bug — the INPUT SCALE is. Against a fixed 5% bar the
RAW edge already clears A on 100% of MLB doubles, 89.6% of hits, before any
shading. MLB median raw edge is 60%, twelve times the bar. Decisive test:
apply the sharpest score in the spec (f=0.647) to EVERY row — the maximum
the design permits — and 75.8% still clear A (MLB 91.7%). Since f is bounded
<= 1, no achievable shading can close a 12x overshoot. Moving the multiply
from the threshold to the edge does not change the outcome.
This is edge_pct behaving as the 2026-07-29 diagnosis described: a price-free
(proj-line)/line gap whose scale is a function of line size. It is not a
betting edge, so no fixed betting-edge bar is meaningful against it.
2.6 efficient-market over-suppression: CANNOT DETERMINE — zero live rows are
shaded, so there is no efficient market in the data to over-suppress.
Phase 3: takeable tagging was completed in the previous order (migration 034,
1246/1254 rows) and is not repeated. The model-version boundary is again NOT
applied: nothing promoted, so no boundary exists.
Unblocking needs the input replaced, not the multiply moved: p_win vs
fair_prob (both already computed) instead of edge_pct, plus scores FIT from our
own record for the markets we actually grade.
Floor: 313 suites / 3899 tests green (9 new), web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Champion grade UNCHANGED. Push scoring untouched. Additive tags only — nothing
deleted, nothing re-settled.
PART A — THE EFFICIENCY CHALLENGER: BLOCKED, NOT BUILT.
Review Zero came back ABSENT on all three inputs:
0.1 efficiency scores DO NOT EXIST (zero occurrences of market_efficiency /
marketEfficiency / efficiency_score in src/ or web/src/).
0.2 base thresholds DO NOT EXIST (engine1.js has zero `edge` references — the
grade is not an edge-vs-threshold comparison; grade_thresholds.json holds
PROBABILITY bands).
0.3 the +/-0.05 additive efficiency nudge DOES NOT EXIST. The only 0.05s on
the grade path are featureCache.teammate_absence_bump, a bvp_advantage
cutoff, and p*0.9+0.05 inside probabilityEstimator (the 0.5*0.1 term of
the shrink-toward-0.5). There is no additive scaling to replace.
So a challenger differing from the champion in EXACTLY ONE thing cannot be
constructed: there is no additive scaling to swap, no base threshold to
multiply, and engine1.js has zero `sport` references so market cannot reach the
grade. A threshold must exist first — that is R1 of
specs/full-output-grade-mapping.md, an explicitly held separate order. Shipping
R1+R4 together would make the Phase-3 delta report misleading: the re-letter
would be driven mostly by switching to probability grading while being
presented as the efficiency fix.
0.4 coverage: the spec names 5 scores; the live ledger has 11 markets and only
MLB total_bases maps to one. 9 of 11 have no score, so "all scored markets"
cannot be satisfied without inventing 9 numbers.
PART B — LEDGER TAKEABLE TAGGING: BUILT (the deferred C2).
New src/config/takeableStandard.js: floor on the minus side, UNCAPPED plus.
Deliberately NOT valueEngine.isTakeable (the -160..+200 PROMOTION band) — a
+400 prop is not promotable but IS takeable; a test asserts the two diverge on
the plus side and agree at the floor so they can never quietly merge. Absent
price returns null, never false (Number(null) === 0 would tag a missing price
takeable). The floor is POLICY not derived (C1 could not derive one) and is
labelled so; each row records takeable_floor so a re-derivation can re-tag.
Migration 034 (applied + tracked): ledger_entries.takeable boolean +
takeable_floor numeric, nullable, partial index. Forward tagging in
ledgerService at row build; backfill in one statement.
Result: 1254 rows, 1246 tagged (781 takeable / 465 below floor), 8 NULL with
null_despite_price = 0 (the NULLs are genuinely priceless rows). Settled 1163
and graded 1254 unchanged.
PART C — the model-version boundary tag is DELIBERATELY NOT APPLIED: no scaling
change shipped, so no boundary exists, and stamping one would mark a model
transition that never happened. modelEras.js is its home when a real one lands.
Floor: 312 suites / 3890 tests green (8 new), web build exit 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Report-only. No threshold, grade, or efficiency value changed.
VERDICT: FLAT. marketEfficiency.js does not exist (zero occurrences of
market_efficiency / efficiency_score in src/ or web/src/). The spec's
0.85/0.60/0.55 values appear in grade_thresholds.json only as PROBABILITY
BANDS - a coincidental numeric overlap, not efficiency scores.
The base edge thresholds (MLB A:5%, NBA A:7%) do not exist either: engine1.js
has zero `edge` references, so the live grade is not an edge-vs-threshold
comparison at all. The specced rule threshold = base x efficiency has no host.
DISPOSITIVE: engine1.js contains ZERO `sport` references. computeFactors
receives no sport or market, so per-market OR per-sport scaling is structurally
impossible in the live grader - not merely unwired.
Phase 2: the matched-edge test is confounded (edge is not the grading input -
the same market emits both B and C at one edge). The aggregate that
discriminates: mean grade index wnba points 4.71 at mean edge 10.2 vs mlb hits
4.58 at 69.5 vs mlb total_bases 4.32 at 84.9 - the efficient market earns the
highest grades on one-eighth the edge, the opposite of spec.
PREMISE CORRECTION (measured): this order's opening claim that full-output and
collapsed grades "agree 100%" does not hold - on 512 rows carrying both they
agree 17.8%, with 33.8% differing by 3+ tiers. The prior discrimination result
stands (champion r=0.0050 null vs probability r=0.1313; MLB 0.0686 n.s. vs
0.2356 p~0.0004). Repo unchanged between orders. The collapse was not a
phantom and the re-adjudication list stays open.
Scope: flat thresholds are a grade-CALIBRATION gap only - the projection and
the CLV edge (which measured p_win, never the letter) are untouched, so this is
not a third shadow-model alarm. But the fix is NOT independently bounded: with
no threshold step to multiply, efficiency scaling presupposes probability
grading. It is rule R4 of specs/full-output-grade-mapping.md and belongs to
that MLB-first challenger.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Report-only. Nothing built, reconnected, or promoted.
Premise corrected again: the three-layer engine is BUILT but NOT WIRED and NOT
DEPLOYED (0 python refs in every grade-path file, 0 python in Dockerfile; there
is no engine1Adapter). So no posterior/CI/similarity prior exists to inventory
or diff. Measured against the collapse that actually exists instead.
THREE collapses, not one: (A) estimateProbability's components discarded at
analyzeViaEngine1:521-524; (B) THE SEVERE ONE - p_win never reaches the grade
at all (engine1.js has zero probability references), so the probability is
excluded from grading rather than collapsed into it; (C) grade_thresholds.json
(probability->grade) read backwards to manufacture confidence.
Market-efficiency scaling is never computed - a gap, not a collapse.
MEASURED on 354 settled rows carrying the served letter and the locked pre-game
p_win (forward, not lookahead). Grade->outcome point-biserial r: champion letter
0.0050 (p~0.93, null) vs probability letter 0.1313 (p~0.013). Per sport: MLB
champ 0.0686 n.s. vs prob 0.2356 (p~0.0004); WNBA champ -0.0986 vs prob -0.1258
- BOTH INVERSE. The served letter is inverted between its only two populated
tiers (B 52.4% n=168 vs C 56.9% n=174).
Verdict: costly on MLB, and un-collapsing does NOT help WNBA -> the challenger
must be MLB-FIRST. Five falsifiable mapping rules specced, incl. R2
(uncertainty grades down) stated explicitly and droppable if it fails.
Hard requirement on the next order: persist per-row n, SE and pre-adjustment p,
or R2/R4 can never be adjudicated (not stored today).
Re-adjudication list flagged incl. proj-v1.1's NOT PROVEN verdict (judged
against the collapsed champion, so not final) and ROI-by-grade (with B/C
inverted, the MLB-C +4.57% segment is likely an artifact of a meaningless
letter).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Archaeology only; nothing built, reconnected, or promoted.
The champion is two DISCONNECTED estimates: the letter is engine1's additive
factor index (zero references to p_win or any probability in engine1.js), and
p_win is probabilityEstimator's frequencyOver + 5 heuristic layers, computed
after and merely attached. The live grade path never calls the Python service.
The Python three-layer engine is NOT DEPLOYED — no python/pip in the
Dockerfile; app.js only health-checks it. So Layers 1-2 never shipped.
Layer 3 is wired BACKWARDS: grade_thresholds.json maps PROBABILITY->GRADE and
the live JS reads it in reverse to manufacture confidence from an
already-chosen letter. Per-sport market-efficiency scaling is specced-absent.
Consequence stated plainly: every metric audited to date is on the shadow
model, not the specced engine, which has never been measured.
Sport boundary TESTED not asserted: a new sport on the live path is a ~10-file
core edit with four documented silent-failure modes. Per-sport records DO
exist (sports.mlb n=526/62% vs pooled overall n=937/58%, each n>=20 gated),
but /api/accuracy ignores ?sport= and the pooled overall would absorb a new
sport. Park x weather confirmed challenger-only; xwOBA and leash absent.
Recovery map is dependency-ordered with MLB as the reference module.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
MLB decided overs n=296, 100% with locked_odds. ROI by locked-price bucket
shows every 95% CI containing zero; the curve is NON-MONOTONE and runs opposite
to the premise (deepest buckets positive, the -111..-160 middle most negative);
and price bucket is confounded with market (+200up = doubles/HR longshots).
Rows needed per bucket to resolve a 5-pt edge: 661-2285 vs actual 8-71 (~187
days for one bucket at current accrual). The inherited -160 is neither
confirmed nor refuted. The no-ceiling call is not supported by this data either
(+200up is the worst bucket) though it is not refuted - it stays a design
choice, not a data-backed one.
Recommends C2 proceed with -160 as an explicitly-labelled POLICY floor plus a
re-derivation trigger (any negative bucket n>=300, or end of MLB regular
season; adopt a derived floor only when a bucket CI excludes zero). Enumerates
all 9 takeable sites, incl. the live drift hazard (backend env-tunable,
frontend hardcoded) and the user-visible band copy in PriceTriplet.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
The anonymous live order (Brionna Jones edge 29.4 ahead of Rhyne Howard edge
42.9) is only explicable by server-side p_win ranking (.90 vs .745, both
takeable) while the payload carries no paid fields — the free caller got the
paid RANKING without the paid SIGNAL, on live data.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
New READ endpoint. No grade, ledger row, lock_line, or scoring write. Push
scoring untouched.
REVIEW ZERO CORRECTED THE PREMISE: the handler NEVER EXISTED in any commit
(searched git rev-list --all for a /top-graded definition in src/ — zero hits).
Not "removed" — the three axios callers (cheatsheetGenerator, gradeOfTheDay,
widget) and the Next proxy were written against a phantom endpoint, so those
three content generators have silently received [] for their entire life.
Contract recovered from the four consumers, not guessed: {props:[...]},
?sport=UPPERCASE (absent = all sports, which gradeOfTheDay relies on) + ?limit,
rows carrying player/stat/line/direction/sport/grade/confidence? plus the
player_name/stat_type aliases and game_id.
POPULATED-PATH RISK FOUND: the board's populated branch had never run in prod,
and dashboard/page.tsx:463 calls g.stat.replace(/_/g,' ') UNGUARDED (g.player
also feeds the row key, /scan URL and heading; sport must be UPPERCASE for
SportPill). toRow requires non-empty string player+stat and a finite line,
uppercases sport, and DROPS unrenderable rows — a shorter board beats a broken
one.
THE LEAK BOUNDARY (why this is server-side): the browser cannot rank on p_win
for all tiers because stripModelPrice deliberately withholds it from unentitled
tiers. Order of operations is
read cache -> RANK with p_win (every tier) -> map rows incl. model fields
-> stripModelPrice(rows, tier) -> serialize
so a free caller receives the paid RANKING without the paid VALUES. Tier comes
from resolveTierFromRequest, which FAILS CLOSED to 'free'. Cache-Control is
private under a bearer token, public otherwise (the /api/snapshot precedent).
ONE SHARED DEFINITION, no drift: new src/utils/gradeRanking.js
(takeablePWin/descNullsLast/rankGrades). heroPropService now imports
takeablePWin instead of its inline copy (behaviour unchanged — it was that
logic verbatim); the selector imports rankGrades; web/src/lib/slateAdapter
keeps its mirror (the browser cannot import src/, S25) and a test cross-checks
the two on identical fixtures (playerName.js precedent). Board is grade-first
("top GRADES"), hero is p_win-first ("top read") — they differ BY DESIGN and
agree within the leading tier.
HONEST LIMIT: the Next proxy (cachedBackendJson) sends no Authorization header
and caches under a shared key, so via the dashboard every viewer gets the
free-tier payload — correct order, no paid values. That is the SAFE behaviour;
forwarding auth into a shared cache is exactly how a paid payload leaks to
anonymous viewers. Per-tier delivery through the proxy needs a tier-keyed cache
and is not done here.
Verified on real prod snapshot data (anonymous path): MLB 8 props, WNBA 10,
0 paid-field leaks, render-contract safe on every row, sport uppercase.
Floor: 311 suites / 3882 tests green (18 new — leak test uses POPULATED p_win,
not today's nulls: entitled gets p_win and it drove the order, unentitled gets
a byte-identical order with all five MODEL_FIELDS absent and no trace in
JSON.stringify, while book/fair market facts survive). Web build exit 0.
Dashboard visual is auth-gated -> tagged for the Chrome audit, not faked.
Held: edge_pct rescale/retirement (Order B); board columns/contract unchanged;
tier-keyed proxy caching.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Display ORDERING only. No grade, ledger row, lock_line, scoring, or edge_pct
scale/display change. Push scoring untouched.
Two defects removed from selectTopGrades (wrong at ANY scale, independent of
edge_pct's separate retirement):
1. edge: Math.abs(numOr(g.edge, -Infinity)) — abs() on an already-
direction-signed value ranked the model's strongest DISAGREEMENTS level
with its strongest agreements (177 public ledger rows carry a negative
edge; positive = the model AGREES with the graded side).
2. Math.abs(-Infinity) === Infinity, so a row with NO edge sorted FIRST —
absent data presented as the top pick (the Number(null) class).
New key: grade -> confidence -> takeable-gated p_win (nulls LAST) -> SIGNED
edge (nulls LAST) -> input order. Scales are never mixed in one comparator.
Takeable band = web valueState.isTakeable, asserted byte-equal to the hero's
config/valueEngine.isTakeable (-160..+200) incl. strict-null.
Alt-line ladder (analyzeViaEngine1:506) no longer sorts by edge_pct: ordered
highest-p_win-first derived analytically at zero added compute — P(stat >= k)
is monotone non-increasing in k, so p_win-desc is line-ASC for an over and
line-DESC for an under. base stays marked; no consumer depends on
alt_lines[0]; deskShowcaseService.rungsOf already re-sorted by line.
THREE PREMISE BREAKS found report-first, before code:
- /api/props/top-graded 404s in prod (absent from src/) so the dashboard
board renders receipts/empty — the edge sort orders nothing there today.
The prior order's "97.3% of rows tie" was a LEDGER measurement wrongly
extrapolated to that board. Fix is correct-in-itself and lands when the
feed is restored.
- p_win cannot be a client-side key for all tiers: snapshotGating strips it
for unentitled tiers ("shipping p_win is shipping the model price").
Verified live: prod /api/snapshot carries p_win on 0/8 MLB, 0/25 WNBA.
- Ladder rungs carry no per-rung price, so the takeable gate is inapplicable.
Verified on real data, both sports, both paths: unentitled — WNBA (n=25)
ordering CHANGED, MLB (n=8) unchanged, signed edge non-increasing in every
(grade,confidence) tie group (20 pairs, 0 violations); entitled — 40 real
ledger rows with p_win+locked_odds, p_win-descending, untakeable chalk NOT
promoted (Trea Turner .757 @-275 does not beat Rhyne Howard .745 @-120)
(36 pairs, 0 violations).
Hero consistency, stated honestly: same signal + same gate, different
precedence BY CONTRACT (board = grade-tier-first "top GRADES"; hero =
p_win-first "top read"). Identical within the leading tier (verified); across
tiers the board may lead with an A the hero doesn't pick. Not a contradiction.
Floor: 310 suites / 3864 tests green, web build exit 0. Dashboard + Desk
visuals are auth/feed-gated -> tagged for the Chrome audit, no visual faked.
Held: edge_pct rescale/display retirement (Order B); building the missing
/api/props/top-graded selector; exposing p_win to unentitled tiers.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
Review Zero found the hero's ACTUAL behavior was worse than "unknown": it ranks
on ev_pct (heroPropService v2), but ev_pct is NULL on served grades and
Number(null)===0 made Number.isFinite(Number(null)) TRUE — so every prop tied at
EV 0 and the "top read" was really the FIRST takeable A/B prop in cache order
(arbitrary, dressed as ranked).
v3: rank by the CHAMPION's p_win (the only signal with a promising, not proven,
edge — its takeable-MLB-over CLV survived the skew audit) among A/B, TAKEABLE-
priced reads (isTakeable band -160..+200, same as the proof/audit). Strict
null guard kills the Number(null)=0 bug. Takeable filter is mandatory (raw p_win
crowns -300 chalk). NO backfill: nothing qualifies → honest empty state
(available:false, reason:'no_qualifying_read'), never a weak recent read.
p_win is RANKING-ONLY, server-side — toHero never exposes it and the route strips
it. Framing unchanged in substance (model number vs book number, grade,
timestamp) — no proven-edge / +EV / best-bet claim, no CLV/ROI/edge number.
Display-only: reads snapshot caches, writes to nothing (no grade/ledger/lock_lines).
Full suite 3852 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
BookComparison.tsx was built but UNROUTED (dead). Route it to the GradeResultCard
via a new self-fetching BookComparisonPanel that reads the live /api/books feed
(source:'bookprices' — the snapshot-locked, fenced, byte-identical store).
Contract fix (Review Zero 0.1): books frequently sit at DIFFERENT lines (WNBA DK
21.5 / FD 18.5; MLB 2/3), so BookComparison now renders EACH book's own line
per-row — never one shared header line implying a false same-number comparison.
Honest states: single-book (the common case for MLB) → one book, "One book
posting this prop.", NO crown/second row; multi-book → all books' own line+price,
NONE crowned (BOOK_CROWN_ENABLED=false — no best-price claim, verified live
crowned:false); no books → renders NULL (panel self-hides), never a placeholder.
No regression: only the always-empty inline d.books section was replaced; grade,
projection, PropLine line, and PriceTriplet price are untouched (wiring test
asserts them). Freshness (0.4): bookprices is written in the SAME snapshot that
locks the grade (intraday refresh touches neither gradedAt.line nor bookprices) —
same fresh, no stale-label needed. Web-only → grade byte-identical trivially.
Full suite 3851 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
The over-side skew audit's confirming check — was our locked line stale-high vs
consensus AT LOCK — was BLOCKED because multi-book lines at lock were never
persisted (bookprices is Redis current-only). This persists them.
- migration 033: lock_lines table (tracked + applied to prod). One row per
(graded prop × book) with both odds + a lock timestamp. RLS enabled, NO
policies -> service-role only (fence). UNIQUE key -> idempotent re-runs.
- lockLineCapture.js: buildLockRows (pure, graded-props only, honest-absent
single-book) + idempotent upsert persist. Built from the in-memory props at
the LOCK moment (ts) -> no Redis re-read, no TTL race.
- snapshotService: persist right after `enriched` (the lock moment; gradedAt
uses the same ts). Best-effort + fenced.
FENCE (measurement-only): lock_lines is read by NOTHING on the grade path
(gradeSlateService, snapshot dedup/indexOdds, challengers, selector, ledger) —
a grep test asserts it, and RLS locks it to the service role. Grade byte-
identical proven: runSnapshot grades are identical with persist on/off (test).
Volume ~1.5-3k rows/day (graded props x books x 5 snapshots); weeks retained,
no pruning needed short-term. Does NOT retroactively fix the existing 62 rows —
future accrual only; confirmation still needs weeks of settled rows. Full suite
3842 green, web build exit 0. No grade/locked_odds/outcome/served surface changed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
Read-only proof. N-gate passed (overlap 45). proj-v1.1 edge-CLV partial
correlation controlling for price = 0.245 (n.s.); ~half the raw 0.455 is the
shared -fair_prob_lock term (mechanical). Champion out-predicts proj on the
same rows (champ partial-CLV 0.380 sig; champ-edge->hit 0.25 vs 0.12). Unders
contaminated (CLV -9.3); WNBA proj-v1.1 doesn't run. Promotion HELD; the
under-audit is moot since proj loses to the champion first.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
closing_prob 59 -> 406 (MLB 248, WNBA 158). Root cause was attachClosingProb's
truncated read + write-once market_unavailable, not capture or the join. CLV
measured: MLB unders lag the close (mean -9.1 prob-pts, 74% lose), MLB overs
+2.0, WNBA flat -> the +4.57% MLB-C and over/under asymmetry are substantially
stale-line artifacts. Unblocks the proj-v1.1 proof order.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
The closing_prob funnel collapsed 100k priced captures -> 59 usable. Root cause
(VERIFIED against prod, join key is PERFECT with 0 mismatches):
- attachClosingProb read closing_captures with .limit(50000) and NO ORDER BY on
a 730k-row table that is 86% refusal rows -> saw ~7% for MLB, missed most
priced closes and declared 200+ rows closeless that HAD a capture.
- market_unavailable_reason was write-once/terminal, so a row wrongly declared
(truncated read / premature declaration before the capture was visible) could
never recover even once its genuine capture existed. 298 rows (204 MLB + 94
WNBA) were stuck this way.
Fix (CLV computation only — no grade/locked_odds/outcome touched):
- Read ONLY priced captures (missed_reason IS NULL, both odds NOT NULL), scoped
to the candidate rows' game_dates -> small AND complete, no arbitrary truncation.
- Drop the market_unavailable exclusion from candidates; make it a re-checkable
absence: a genuine close now UPGRADES the row (writes closing_prob, clears the
verdict). closing_prob stays write-once (first true close wins). No capture +
past game -> still declared absent (honest). No churn on already-absent rows.
- New internal trigger POST /api/internal/ledger/attach-closing[/:sport] for
backfill + verification (scheduler already runs attach per tick).
Recovers ~312 usable closes (59 -> ~371), MLB included. Capture itself was
healthy all along (94.9% MLB / 95.8% WNBA per-prop coverage). Full suite 3835
green (17/17 instrument tests incl. 2 new recovery cases), web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
NexaPay was cross-project contamination (from another venture) — never a real
VYNDR payment path. Purged; Stripe path untouched.
Removed:
- web/src/services/nexapay.ts (createPaymentLink/getTransaction/HMAC verify)
- web/src/app/api/webhook/nexapay/route.ts (the only importer; Next-registered,
reachable — now gone)
- NexaPay comments in email.ts + checkout/route.ts
- Active NexaPay entries in docs/SYSTEM-MANIFEST.md (route list, NEXAPAY_* env
table, service row) + stale claim in wiring-data-train.md
- sw.js precache entry for the deleted webhook chunk
Verified: ZERO NexaPay in code (web/src, src, tests). Full suite 3833 green
(count unchanged — nothing depended on it, confirming it was dead). Web build
exit 0. sw.js parses clean. Stripe checkout untouched (Next→Express→Stripe).
FLAGGED FOR KEV (a repo delete cannot close these):
- Coolify env: remove NEXAPAY_API_KEY / NEXAPAY_WEBHOOK_SECRET / NEXAPAY_API_URL
- Revoke the NexaPay API key + webhook secret at NexaPay's dashboard; de-register
the webhook if an account was ever configured
- DB column user_profiles.nexapay_customer_id is orphaned (no reader/writer) —
drop via a follow-up migration (migration 011 left as history)
Cross-project check: ZERO Noctem-Supabase refs; VYNDR references only its own
Supabase (zmdnczhtdxcddsxzttub). NexaPay was the sole contamination found.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
Records the six live fabrications removed/hidden, keeps media/newsletter/WIRE on
the board as real work, logs the news/line-movement signal as a future model
input, and logs the known honesty gaps (hit-rate-without-ROI, CLV starved).
Honest state: "no KNOWN live fabrications," not "provably none."
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
Six live untruths corrected — no grade/snapshot/scorer/pipeline touched:
1. /compare — hardcoded Jokic A+/Wembanyama A + fake VERDICT replaced with an
honest in-development state; removed from Nav + BottomTabBar (route still
resolves, never the sample). Real two-player build is later.
2. Pricing — founder Desk $34.99→$44.99 (matches lib/checkout.js), Analyst
$14.99; removed the struck $19.99/$44.99 "regular" numbers and DeskShowcase's
stale $34.99. First-100 counter is real (ClaimMeter → Stripe countFounderSeats);
no fake "first 50" desk claim added (no such counter exists).
3. FAQ "NexaPay" → Stripe (verified: live checkout is Next→Express→checkout.stripe.com).
4. FAQ + Features "Brier/CLV published from day one" removed (not surfaced yet) —
returns when real. Backend Brier compute untouched.
5. MobileEdgeBoard removed from the Slate — its edge% feed was a miscalibrated
placeholder (masked >40% as "—"); phones now show the real game cards.
6. Price triplet — never-computed model/EV now derives NO_MODEL (honest absent,
MODEL "—" / "NOT PRICED", no verdict) instead of QUARANTINE's false "we
suppressed our price / a leg is poisoned" copy. Fixes grade card + LiveHeroProp.
Full suite 3833 green, web build exit 0. Tests updated to the new honest contracts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
Read-only inventory order — no code built/wired/fixed/deployed. Maps every
user-facing surface and model component against DESIGNED·BUILT·WIRED·LIVE·HONEST,
re-derived from repo b0a51c8 + prod + design bundle (not STATE.md narrative).
15/26 surfaces fully done. Names the graveyard (BookComparison, ShareCard,
proj_ladder, arch-v1/contact-v1 ledger-only, /intelligence orphan) and the live
honesty gaps (/compare hardcoded, FAQ NexaPay/Brier, MobileEdgeBoard placeholder,
EV fields NULL on served grades). STATE.md now points to the matrix as canonical.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
Per-book prices existed only transiently (odds cache, ~1h, raw names, grade-path
input); every grade-path persistence point collapses to one book. The
/api/books feature was built+mounted but non-functional (fed FLAT rows to a
GROUPED comparator -> always empty).
Phase 1: bookPriceStore captures per-book prices from `props` BEFORE dedupeProps,
keyed nameKey|stat, into bookprices:{sport} (SNAP_TTL) in snapshotService. Fenced:
reads props, writes its own key, read by nothing on the grade path. Grade proven
byte-identical (test + no-grade-path-reference grep test).
Phase 2: scripts/measure-book-spread.js reports same-line best-vs-worst spread
(cents + implied-prob pts), per sport, never pooled. Pre-registered crown
threshold: median >=8c OR >=2pp. Runs post-deploy on real data.
Phase 3 (backend): compareProp is honest-absent (single-book/flat -> no crown)
and the crown is gated (BOOK_CROWN_ENABLED, default OFF until Phase 2 clears).
/api/books repointed to the snapshot-locked store (fallback odds cache),
nameKey-matched; `source` field is the deploy fingerprint.
HELD unchanged: dedupeProps, snapshot dedup, selector, grade, champion,
challengers, ranking, edge_pct/ev_pct. UI routing of BookComparison + crown
treatment deferred to post-measurement (gated on Phase 2). Full suite 3834 green,
web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
proj_book_implied derived from raw book_odds — VIG-INCLUSIVE. A -110/-110 market
implies 52.4%/side (104.8% sum); fair is 50%. Comparing our P against raw book
overstates the book on both sides, biasing the handicapper test IN OUR FAVOR; on
juiced longshots (the Judge HR -18.5pt case) much of that "edge" was vig, not
disagreement.
Fix (fenced to proj-v1's stored comparison basis): proj_book_implied now derives
from DE-VIGGED FAIR via the grade's g.fair_prob — the SAME multiplicative de-vig
the triplet uses (utils/devig.js), so the basis matches the product's shown fair.
Expressed on the OVER basis (under props → 1 - fair) to match our stored P(≥rung);
traded-rung ladder book_implied likewise. HONEST-NULL where fair is uncomputable
(one-sided market, ~14%) — NEVER a raw-book fallback (that would recreate the vig
bias on a subset and mix two bases in one ledger). proj_factors records
book_implied_basis ('fair_multiplicative'|'none').
Phase 0 (prod-verified): fair reachable at store point (g.fair_prob on the grade,
no threading); 86% batting coverage; method = multiplicative/proportional.
Phase 2 FLAG: multiplicative de-vig mis-splits vig on juiced longshots (favorite-
longshot bias), so a longshot fair still carries known method bias — flagged
per-row (longshot_devig_caveat); a better de-vig (Shin/power) is a separate item.
Phase 3: version bumped proj-v1 → proj-v1.1 so pre-fix (raw-book) and post-fix
(fair) rows never silently mix — the projection model is byte-identical, only the
basis changed; pre-fix rows can't be recomputed (only the graded side's odds were
stored). Champion + arch-v1 + contact-v1 + proj-v1's other columns untouched.
proj suites 26/26.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The live fingerprint showed proj_book_implied null on every real row: grades
carry book_odds/locked_odds (e.g. -264) but NOT a de-vigged fair_prob, so keying
the book comparison off fair_prob yielded null. The book ODDS are exactly "the
book's implied probability" the handicapper test needs. Now proj_book_implied +
the traded rung's book_implied derive from americanToImplied(book_odds),
expressed on the OVER basis (under props → 1 - implied) so it's directly
comparable to our P(≥rung). Vigged (a known offset the ledger measures both
sides of). proj-v1 suites 24/24.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
1. matchupRead fly-ball signal: the batter metrics `gb_pct_bb`/`fb_ld_pct` are
MISLABELED — they're exit velocities by batted-ball type (Judge fb_ld_pct =
100.3 mph, not a rate), not ground/fly RATES. Switched fly-ball lean to
avg_launch_angle (league p10/p50/p90 = 7.1/13.9/20.1°), the correct signal.
2. Absolute rate now fits the FULL season (recency-weighted), not a 20-game
window: the window under-sampled rare stats — Judge HR projected 0.11 vs his
0.28 season rate (a fake -32pt edge). Now point=0.27 (matches season); the
last-5-2x recency lean is preserved.
Post-fix induction (real statsapi logs + real statcast): Judge HR 0.27 (P>=1
0.235 vs book 0.42 -> flags the juiced over), Judge TB P>=2 0.548 vs 0.48
(+6.8pt), thin-hot 3-game P>=1 0.726 / P>=3 0.164 (credible low, thin high),
.300 hitter != 3.0. proj-v1 suites 23/23.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
A THIRD challenger (after arch-v1, contact-v1), MLB batting v1. Champion is
market-relative P(stat>LINE); proj-v1 is ABSOLUTE — what the hitter will DO —
emitted as a full distribution from which the WHOLE LADDER (P≥1,P≥2,P≥3) derives.
Champion untouched; nothing claimed; the ledger decides per rung, per stat.
- projection/distribution.js — Bayesian Gamma-Poisson → negative-binomial
predictive. Admits over-dispersion; under-dispersion → Poisson approx
(conservative, documented). Uncertainty scales with sample by construction
(r=α): thin → WIDE (real mass on P≥1, honestly thin P≥3), thick → tight.
NEVER abstains — width carries the honesty.
- projection/matchupRead.js — the input the book doesn't use. HONEST FIDELITY:
pitcher repertoire is rich (97% pitch-mix) but hitters have NO pitch-type
performance, so TRUE repertoire-vs-profile is impossible today. This is the
COARSE version (arsenal buckets fastball/sinker/breaking + whiff/hard-hit
tendency × hitter whiff/chase/gb-fb/hard-hit) — beats generic L/R, derived +
documented + TESTED two-sided. A hitter pitch-type feed unlocks the true form.
- projectionChallenger.js — park RELATIVE to the player's own log exposure
(isHome→own park, away→opp park; Phase B's raw-multiply bug solved), recency-
weighted fit, per-factor breakdown (form/park/weather/platoon/matchup — show
your work), full rung set + book-implied per rung. Combined non-form
multiplier bounded.
- Wired after contact-v1, own try, flag PROJ_V1_ENABLED, reusing arch-v1's
already-computed park/weather/platoon (no duplicate env I/O). Own ledger
columns (migration 032, applied to prod): distribution, ladder, point, line,
our-P, book-implied, factor breakdown — measurable per rung/stat after settle.
Phase 0 (prod-verified): venue join via isHome; NB family; uncertainty-as-width;
coarse matchup honest fidelity; no lineup-slot (per-game rate, volume implicit).
Sanity: thin-hot → wide (credible low rung, thin high rung); .300 hitter ≠ 3.0;
matchup two-sided; champion byte-identical. proj-v1 suites 23/23; snapshot/
ledger/siblings 74 green. Forward-only, version-stamped, PROJ_V1_ENABLED kill.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Phase A #2: the champion grade (l5/l20 result-based form) is a HYPOTHESIS that
contact quality predicts better — unmeasured on our props, with zero settled
p_win yet. Swapping l5/l20 (the champion's two heaviest ±1.0 factors) blind
could degrade the core grade undetectably for weeks. So this NOMINATES contact
quality as a second challenger, records what it WOULD project per prop, and lets
the settled ledger decide. Nothing users see changes; the champion is untouched.
- src/services/contactChallenger.js — pure, mirrors challengerProjection. Log-
odds lean (capped, never a re-forecast) from SEASON contact quality vs league
percentiles. Metric→prop mapping is the whole game: barrel_pct→HR,
hard_hit_pct→TB/doubles, k_pct-INVERSE→hits (singles resolve on contact
frequency, not barrels), k_pct→batter K. rbi/runs/walks ABSTAIN (opportunity/
discipline — no clean contact predictor). Honest-absent: thin (<50 PA)/absent/
unmapped/non-batter → p_win_contact NULL (no projection), never a fallback;
"measured but unremarkable" is distinct (equals champion, delta 0).
- Wired in snapshotService AFTER arch-v1, reusing the already-loaded statcast
rows; its own try so a second challenger can't break the pipeline. Reads
g.p_win, never writes it.
- Retained SEPARATELY on the ledger (p_win_contact/contact_delta/
contact_adjustments/contact_version='contact-v1') so each challenger's marginal
contribution is measured independently; ledger_entries.stat gives per-prop-type
segmentation. Migration 031 (applied to prod).
Phase 0 (prod-verified): statcast_aggregates is SEASON cumulative (not rolling),
48h stale now but season-scoped so ~8 PA/600 is negligible; 100% of graded
hitters covered, 92% at ≥50 PA; no xBA/xwOBA in the feed. Forward-only,
version-stamped (contact_version null on pre-nomination rows). Promotion is a
LATER decision on settled evidence, per prop type — never asserted here.
contactChallenger 14/14; snapshot/ledger/arch-v1 suites 80 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Semantic COLOR fix, not cosmetic. The S6/S7 build predated the current Scanner
States spec: it used AMBER (the quarantine / model-suppressed channel) for the
"no market / line not priced" case, telling users "model suppressed" when the
truth is "the board never priced this." Corrected to the HANDOFF Session-3
blue-boundary law: BLUE (--priced-out #8FB2DE) = no-market boundary; amber stays
QUARANTINE; red stays REFUSAL.
Phase 1 (reskin): NoMarketState → dashed BLUE void box + blue header/copy; the
S7 rows → spec format (o 27.5 · BK −114 · ◆ −105 · OPEN READ ▸) at 44px,
390-legible. Input-area surfacer pills reskinned to the blue channel too.
Phase 2 (never-built states, only those Phase 0 confirmed against live data):
- GREEN CTA with LIVE player count ("PLAYER · N PRICED PROPS ▸"), degrading
honestly to the board path ("N PROPS LIVE · TONIGHT'S BOARD ▸") at 0 — count
from the SAME fresh index as the rows (pricedCountForPlayer), can't disagree.
- CASE A none-priced DEFAULT: "WE PRICE THESE FOR [player]" — the player's other
priced stats (pricedStatsForPlayer, filter by nameKey).
- Typed-line-mismatch blue fact line ("o X ISN'T PRICED · NEAREST ↓").
FLAGGED / not built (no shells): Case C off-slate quiet-stop needs schedule/
roster membership the pricedLines index doesn't carry (out of the presentation
fence). Spec CONTRADICTION: Case B says "fair previews amber," but the law
reserves amber for quarantine — fair renders NEUTRAL ◆ (blue-dim), not amber, to
avoid blurring the channel.
Free-tier gate VERIFIED before rendering FAIR: fair_odds is the de-vigged MARKET
price (valueState: "never hide the honest fair number"), NOT the gated
model_odds — no paid leak. Carried through indexPricedLines (additive; keying/
refresh/onPick/stale-tap all unchanged — the proven S7 data path is untouched).
Scan A byte-identical; PRICED_NUDGE_ENABLED still the kill switch; reversible.
Build exit 0; priced + parity suites green (67).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The A/D investigation found CV (std/mean) is scale-broken on count data —
for a Poisson-ish stat cv ≈ 1/sqrt(mean), so EVERY stat with mean < 4 blew
past the boom_bust cutoff regardless of behavior. The S63 stopgap made those
return 'unknown', which silently ate a real +1.0 consistency signal on every
MLB batting prop — steady low-mean hitters never got their earned factor.
Fix, fenced to the low-mean branch of consistencyScore (the only branch that
was returning 'unknown'): classify with the index of dispersion (variance/mean,
Poisson baseline 1.0) — the scale-appropriate, UNBIASED statistic for counts.
mean ≥ 4 keeps the NBA-calibrated CV path BYTE-IDENTICAL (zero NBA blast
radius). This is a bug CORRECTION, not threshold loosening: the CV thresholds
and the engine1 ±1.0 delta are unchanged.
Bands (asymmetric around Poisson 1.0, since counts are naturally mildly
over-dispersed): iod<0.60 elite / <0.85 reliable (+1.0) / ≤1.30 volatile
(neutral) / >1.30 boom_bust (−1.0). Sample floor MIN_GAMES_FOR_IOD=8 so a
thin sample abstains ('unknown') — no small-sample guess.
Validated on real 10-game logs (two-sided): Kwan hits 0.67 / Alonso hits
0.78 → reliable (RECOVERED); Alonso TB 2.57 / Henderson hits 1.33 → boom_bust
(no false consistency); HR mean 0.1 → 1.0 → neutral. Direct engine1 proof: a
strong steady prop that grades B+ today reaches A- once the +1.0 fires; a
boom-bust bat stays B (no inflation). A- now emerges NATURALLY from a real
recovered factor. Standing two-sided test pins all three directions.
Forward-only (settled grades are locked in the ledger, never re-graded).
Emitting A- ≠ proving A- — the A-tier record accrues from emission, still
measurement-gated. Full unit suite green (4 pre-existing redis/timing flakes
pass in isolation); web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Migrate the Landing hero (Landing-only component; LiveHeroProp/triplet
untouched) to the design bundle's visual language: mono "SPORTS
INTELLIGENCE TERMINAL" eyebrow, tightened 56px/800 headline "Every line,
graded before you bet it.", dual mono CTAs (OPEN THE TERMINAL / SEE THE
PUBLIC RECORD), proof chips.
Claims audited against the held list — the bundle's marketing copy carries
claims we have not earned; those did NOT ship:
- "CLV-VERIFIED" chip + "verified against closing lines" — HELD (C4 broken)
- "312 props / 47 games" count — fabricated demo → ABSENT (no count)
Shipped copy is earned only: pre-graded, edge-ranked, settled in public,
misses included; PUBLIC LEDGER / MISSES INCLUDED / 5 FREE READS chips.
HELD + FLAGGED for a design pass (fabrication-backed, no honest designed
state — check-don't-freelance): the edge-board demo (A+/A don't emit), the
"grades calibrate · A+ hit 80%" claim (calibration unmeasured), and the
tier-record band (real data is B/C-only, no ROI). Not built with fake data.
Tokens via var(--g-a/--void/--border/…) with bundle-hex fallbacks
(documented-intentional pattern). Build exit 0; parity QA green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Reference material only, ZERO product code. Adds Vyndr Scanner States.dc.html
(S6/S7 spec) + HANDOFF Session 3 blue-boundary-channel law (#8FB2DE = the honesty
channel: priced-out / no-market / line-not-priced). This is the diff baseline for
the design-migration arc.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
REPORT-FIRST correction. This order's premise — "S7 is a shell that doesn't
update per selection" — is not what the code does. pricedForSelection is a
useMemo on [pricedIndex, selectedPlayer, stat] and setSelectedPlayer/setStat
fire on every user pick, so the chips already update per selection, and the
prior session's verification of that stands. The genuine gap was FRESHNESS: the
snapshot fetch depended on [sport] only, so pricedIndex was fetched once per
sport-change and never refreshed. The pricing cron re-prices at five UTC hours,
so a scanner left open across a cron boundary surfaced hour-stale priced lines.
That is the real defect, and the only one fixed.
FRESHNESS. The fetch is now a refreshPriced callback re-run when the held
snapshot is older than PRICED_STALE_MS (30s, matching the /api/snapshot cache)
at the moment of use — on selection change and on window focus — so a long-open
page never shows a stale line. Sport change still clears the index first, so the
old sport's lines never flash.
STALE-TAP was already safe and is unchanged: the scan submit re-fetches the live
snapshot server-side, so a chip that's gone stale between render and tap either
lands on a real triplet (still priced) or degrades to the honest empty state
(rotated away) — proven in the prior session and re-confirmed here (an
off-snapshot line returns no market and shows the empty state).
REVERSIBLE GATE. The whole nudge sits behind one PRICED_NUDGE_ENABLED flag: false
empties the surfaced set, so the scanner falls back to S6's link-only empty
state with the chips gone. Shipping enabled only after the cases are proven this
session; the flag is the instant revert lever.
DISPLAY-LAYER ONLY. Only scan/page.tsx changed. GradeResultCard, PriceTriplet,
gradeAdapter, valueState, both scan routes and the pure pricedLines helper are
byte-identical — Scan A and the scan-submit resolution are untouched, and the
change is independently revertible.
Tests 3765 passed / 303 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The Session-78 diagnosis stands: the join works, and a marketless scan rightly
shows no triplet. This makes that absence legible and points the user at what IS
priced, without fabricating a market.
PHASE 0 gate — design-check, reachability, timing, all clear. Design-check: the
bundle has the triplet's own REFUSAL language ("we'd rather show nothing than a
number we can't stand behind") as the honesty precedent, and a designed
EmptyState component whose actions give a path forward — so the empty state is
built in the established visual language, not freelanced. Reachability: the
scanner already fetches games/odds/search per selection; the snapshot is one
more public, 30s-cached fetch per sport, re-run when the sport changes.
Staleness: the snapshot rotates 5x/day and every scan re-validates the market
server-side at submit time, so a surfaced line that goes stale degrades to the
empty state on tap rather than a vanishing triplet — the stale-tap guard is
inherent, not bolted on.
Reversibility was the design constraint. The working card, price triplet, grade
adapter, valueState and the scan route are BYTE-IDENTICAL — a test asserts none
of them even reference the new empty state. Everything new lives in two added
files (lib/pricedLines.js, components/vyndr/NoMarketState.tsx) and additive
blocks in the scan page. Removing them leaves the Scan-A path untouched.
Non-fabricating by construction: indexPricedLines only keeps snapshot rows that
carry a real book price, keyed by exact player+stat via nameKey. A different
stat priced for the same player surfaces nothing for the picked stat; an
off-slate player surfaces nothing; nothing is suggested, interpolated, or
rounded to a nearest line. The empty state shows no market numbers of its own —
only real priced lines as one-tap chips, or a link to the live board when there
are none.
Framing is help, not restriction: a "PRICED TONIGHT" chip row sits under the
free-typed line input, and the scanner still accepts any player, stat and line.
Tapping a chip pre-fills the priced line and re-scans it — the market is
re-resolved server-side, so the tap either yields a real triplet or degrades to
the honest empty state.
Path forward, not a wall: a marketless scan no longer dead-ends in blank space.
It states truthfully that the board didn't price that line, keeps the grade, and
routes the user to the priced lines for that exact player+stat or to tonight's
board.
Tests 3760 passed / 303 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Report-only verification of the two unproven claims from the Session-77 wiring,
which fired platoon on synthetic split-less hitters where the K=600 regression
zeroed the multiplier and masked both direction and resolution. No product code
changed — this adds one standing regression test.
FIXTURE — Yordan Alvarez (LHB), real 2026 splits, deep both hands: 115 PA vs LHP
at .529 slg, 327 PA vs RHP at .695, overall .652. The regression leaves a
material multiplier both ways (0.97 vs LHP, 1.023 vs RHP), so unlike the last
test this fixture can actually reveal direction.
RESOLUTION — verified on the live slate that the pitcher-hand attached to a
hitter is the OPPOSING team's probable, not his own. CLE (home) resolved to
Minnesota's away starter 696070; MIN (away) resolved to Cleveland's home starter
800048. The chain — hitter's team, the game, the other team, that team's
probable, that pitcher's hand — is correct, and it is pinned independently of
direction because a backwards resolution is invisible on a neutral hitter.
DIRECTION — deterministic L-vs-R on the frozen Alvarez fixture. Facing RHP nudges
UP (1.023) because he slugs .695 there, above his .652 overall — a favorable
opposite-hand matchup, exactly what platoon theory predicts for a left-handed
bat. Facing LHP nudges DOWN (0.97). The two move opposite directions, and
crucially the SPECIFIC sides are asserted, not merely "opposite" — a
mirrored-but-inverted implementation would put RHP below 1 and fails here. Both-
backwards is ruled out.
The test also pins a REVERSE-split hitter, Brandon Nimmo, who hits better vs LHP
than RHP. His multiplier goes up vs LHP, following his real numbers rather than a
hardcoded LHB-vs-RHP assumption — proof the sign is data-driven, which is the
correct design.
VERDICT: PASS. Resolution correct, direction correctly signed against both the
real split and platoon theory. Pinned by tests/unit/platoonPolarity.test.js so
the polarity cannot silently regress — the opp_rank_stat lesson applied.
Tests 3750 passed / 302 suites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Verified state going in: parkBase, weatherMod and platoonSplits were called by
nothing, and env_multiplier was non-null on zero rows across four orders. The
adjusters were correct in isolation and starved of inputs. This gives them their
inputs and changes none of their internal logic — the five adjuster files are
byte-identical after this commit.
PHASE 0 GATE — all three inputs are available at snapshot build, and the two
join keys already existed. Venue: always, on every schedule game object.
First-pitch: always, gameTime on the same object. Opposing-pitcher hand:
present once the probable is declared, via the pitchers endpoint's pitcherId
joined to statsapi handedness — 15 of 15 games declared this afternoon, though
morning locks precede declaration and those props honest-absent on platoon,
correctly. The batter-handedness join (statcast bats) and the MLBAM id were
already on each grade from earlier sessions.
environmentContext.js is the wiring, kept separate from the adjusters so they
stay pure. It fetches once per snapshot: the schedule (team to venue, gameTime),
probable pitchers (team to opposing pitcher id), one batched handedness call,
one Open-Meteo forecast per home park, and batter splits per graded hitter. Park
coordinates for 30 parks live here as public geometry, the same class as the
dome list and centre-field bearings already in weatherMod, rather than inside an
adjuster. Everything is best-effort: a missing venue drops park and weather, an
undeclared pitcher drops platoon, and any fetch failure degrades that prop to
archetype-only rather than breaking the pipeline the adjusters are measured
inside.
attachChallenger becomes async and takes a per-grade contextFor that returns the
environment coefficient (park_base x weather_mod, composed) and the matchup
(platoon). Point-in-time holds: the weather is a forecast for first pitch fetched
now, and the split is the hitter's line entering the game — neither reads a
settle-time value.
Attribution is independent. env_multiplier, env_park_base, env_weather_mod and
env_weather_state land in their own ledger columns, and challenger_adjustments
keeps every axis — archetype, environment, matchup — as a separate entry, so
when volume accrues each of the four can be measured for its own marginal
contribution rather than as one blended delta.
The combined move stays bounded, tested on the worst case: a Coors slugger with
wind out and a favourable platoon, all at once, still moves under 12 percent,
because every layer is capped and the total nudge is clamped. Stacking leans, it
does not compound into a re-forecast.
Non-MLB honest-absents entirely — park, weather and platoon are MLB-only today,
so a WNBA prop gets no environment and no matchup.
The champion is untouched throughout: p_win is read, never written, the served
snapshot payload is still the enriched object, and a test confirms p_win passes
through byte-for-byte while the challenger moves.
Tests 3741 passed / 301 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The highest-value adjuster and the thinnest sample in baseball. The regression
is not a refinement here, it is the entire feature: applying raw splits would
adjust projections on noise, which is worse than not building it.
PHASE 0 — both gates clear, and one was already closed. Splits are a statsapi
pull, one call per hitter (statSplits with sitCodes vl,vr). The batter-handedness
join that Session 69 recorded as pending is in fact DONE: statcast_aggregates
carries bats for 604 of 604 batters, 210 left, 327 right, 67 switch. STATE said
pending; the data says otherwise, and the note is corrected. Point-in-time holds
as long as the split is fetched before first pitch, since a season split queried
this afternoon cannot contain tonight — but a historical backtest would use
season-final numbers and leak, so clean measurement is forward-accruing.
THE SPINE — regressed = (PA x observed + K x prior) / (PA + K), with K = 600 PA
and the prior being the hitter's OWN blended rate rather than the league's. The
question a platoon adjustment answers is whether he is DIFFERENT against this
hand than he normally is, so his own line is the correct null and a hitter with
no evidence of a split correctly gets nothing. K is deliberately conservative:
platoon skill is famously slow to stabilise, with the half-signal point for
right-handed batters near a thousand PA.
THE MAKE-OR-BREAK TEST, both halves. A .310 average against left-handed pitching
on 30 PA gets 4.8% weight and moves the projection by 0.003 — essentially
nothing, which is the correct answer rather than a limitation. The SAME .310 on
400 PA gets 40% weight and moves it by 0.023, eight times as far. A test asserts
that ratio stays above five, so if the regression ever breaks the suite says so
instead of the projections quietly drifting onto noise.
Real data behaves exactly as the mechanism predicts and is worth recording:
Josh Bell hits .259 against lefties and .248 against righties, which looks like
a platoon split until the sample speaks — 126 PA earns 17% weight and the
adjustment lands at 1.005. Aaron Judge, 76 PA against lefties, comes out at
0.999. Neither is material. Most hitters will get nothing from this adjuster,
and that is the honest output, not a failure.
Honest-absent has five distinct routes, all returning exactly 1.0: no batter
handedness, no pitcher handedness, no splits, a stat platoon says nothing about,
and a missing side falling back to the prior rather than to zero.
INDEPENDENT of the environment. Park and weather compose into one coefficient
because they both describe the stadium; platoon describes this hitter against
this pitcher's hand, so it rides its own slot with its own label. Entangling
them would make both harder to attribute when the instrument scores them.
Directional, mirrored on the under, capped at 15%, and inverted for strikeouts
where a higher rate means a higher prop rather than a better hitter.
Tests 3729 passed / 300 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Completes the coupled environment: effective = park_base x weather_mod. Weather
tilts the park, it never overrides it — a wind-out night at Oracle Park is still
Oracle Park.
PHASE 0 — both feeds are free and keyless. statsapi /venues gives every park's
coordinates in one call; Open-Meteo returns hourly temperature, wind speed and
wind direction for those coordinates hours before first pitch, which is when we
project. Verified live.
THE SPINE — two weather values, two purposes, never crossed. The FORECAST we
held at projection time drives the live adjustment AND is what the instrument
measures, because it is what we actually knew. It lands on the ledger row beside
p_win. The ACTUAL goes only to game_context as raw material for future
self-derived weather factors, and is read by nothing that scores a projection.
Using the actual to measure tonight would be scoring ourselves on information we
did not have. The actual is also pulled from Open-Meteo's ARCHIVE endpoint
rather than the forecast endpoint, because asking a forecaster after the fact
returns a re-forecast, not what happened.
WIND IS PARK-ORIENTATION CONDITIONED. Wind direction is meteorological — the
direction it comes FROM — so blowing out to centre means arriving from the
opposite bearing. Getting that backwards would invert every wind adjustment in
the system, so the 180-degree rotation is commented at the site and pinned by a
test on all three cases: straight out, straight in, and crosswind. Centre-field
bearings are public geometry, in the same class as the dome list; a park missing
from the table gets no wind effect at all rather than a guessed one, and keeps
its temperature effect.
THREE HONEST DO-NOTHING STATES, all multiplier 1.0, none fabricating an effect.
Dome: weather does not apply, and the PARK factor still does — verified that a
domed venue keeps its sub-1.0 park base while weather stands down. Forecast
absent: none available for this park and time. Sub-threshold: a real forecast
below a meaningful bar, because manufacturing a 0.3% nudge on a light breeze is
false precision. Weather also says nothing about a strikeout prop and returns
not-applicable rather than a neutral it might later be tempted to fill.
Conservative and ledger-tunable: every magnitude is an env var, the total is
capped at 12%, and nothing here is asserted. This is a nominated challenger that
earns its place on the instrument or is cut.
Induced at Wrigley, whose centre field bears 32 degrees: wind from 212 at 15 mph
computes as 15 mph straight out, weather 1.12 composed with park 1.06 for an
effective 1.187 and a +0.043 nudge; the under mirrors exactly; the pitcher's
home-runs-allowed prop moves with the hitter's, since both are P(over) on a ball
leaving the park. Wind in drops the coefficient to 0.955. A calm 72-degree
evening, a dome, and a missing forecast all return 1.0 by three different
honest routes, with the park base still applying in each.
One correction to the order worth recording: it describes a wind-out night as
helping the hitter and hurting "the pitcher there's HR-allowed" as opposite
sides. In prop terms both go the same way — the HR-allowed OVER is more likely
too. The sign lives in the stat, exactly as established for park factors, and
the implementation follows that rather than the phrasing.
Migration 036. Tests 3707 passed / 299 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
PHASE 0 — the settle path sees a player's game-log line, not the game. It knows
date and teams, never venue or final totals. But the grain is far cheaper than
per-prop or even per-game: ONE statsapi schedule call per game DATE returns
every game that day with venue, linescore and scoring plays. Fifteen games, one
call, verified live.
PUBLIC BASE — the ingestion was already done. The static FanGraphs table from
Session 15 is the public base; this converts its 100-indexed values into the
multipliers the composable architecture wants (Coors 128 becomes 1.28) rather
than ingesting a second copy of a number we already hold.
It is labelled COMMODITY in the code, not just in a comment. Every resolution
carries a provenance record, and the public one reads proprietary: false with
the note "Commodity: a public number. Not a VYNDR derivation." The proprietary
label exists but belongs only to the self-derived version, and only once it
beats this base on the instrument. A surface rendering a park effect can state
which it is rather than implying the flattering one.
Honest-absent where even the PUBLIC number is thin: a relocated club in a
temporary venue gets no factor, because a public number for a park with one
season behind it is no more trustworthy than ours would be.
SOURCE-PLUGGABLE is the architectural point. resolveParkBase() is the only
accessor, public and derived return identical shapes, and callers never branch
on source — so when self-derived factors clear their floor they swap into the
same slot with nothing downstream to rewrite. A derived source with no factor
available returns absent rather than silently falling back to public, because a
silent fallback would make a proprietary claim out of a commodity number.
GAME-LEVEL CAPTURE starts now because it cannot start retroactively. Game grain,
deduped on game_id, never copied onto prop rows — a game's totals belong to the
game, and duplicating them per prop is how one fact starts disagreeing with
itself. Every field is tied to a named future derivation: venue for park
factors, runs for the run environment, HR totals for HR factors. Nothing else is
stored. Only Final games are captured, since an in-progress total is not a
result, and a game with no scoring plays reports HR as absent rather than zero.
HR totals come from scoring plays, which is complete because every home run
scores at least the batter.
The accrual target is stated rather than promised: 150 home games per venue at
roughly 81 per season means about two seasons before a self-derived factor can
be nominated, and accrualStatus() reports live progress per venue so the wait is
measurable.
Induced: Coors home runs +0.061 for the hitter and identically +0.061 for the
pitcher's home-runs-allowed at the same park, mirrored on the under; San
Francisco negative; Tampa flagged weather-N/A with its factor still applying;
the Athletics' temporary venue absent; strikeouts untouched. A real 2025-07-19
capture produced 15 games across 15 venues, 12 with HR totals, zero duplicate
game ids.
Migration 035. Induce with POST /api/internal/gamectx/:date, progress at
/gamectx/accrual. Tests 3688 passed / 298 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
PHASE 0 GATE — the answer is BOTH, and the important half was already here.
A STATIC FanGraphs park-factor table has existed since Session 15
(src/data/parkFactors.js) and computeFeatures already consumes it, so park is
not a new idea in this codebase. What was missing is OUR derivation. I nearly
built a second source of truth before finding it; the new service lives at
src/services/parkFactors.js and the two are deliberately distinct.
That discovery changes the point of this order rather than just its scope. If
the champion already sees a park factor, adding one to the challenger risks
double-counting — which is exactly the redundancy the Session-72 harness exists
to catch. So park ships as a NOMINATED CHALLENGER whose job is to be tested for
marginal contribution, not as an assumed improvement. Checked and worth noting:
the static table reaches computeFeatures but NOT probabilityEstimator, so it
does not currently touch p_win at all.
DERIVATION, not ingestion. statsapi gives every game with venue, linescore and
scoringPlays in one call per date range — and since every home run scores at
least the batter, HR totals are fully recoverable from scoring plays. Derived
from 5,055 real games across 2022-2025: Coors tops the run environment at
1.099, Dodger Stadium tops home runs at 1.106, Oracle Park and PNC suppress
them at 0.923 and 0.917. Eighteen parks cleared the floor, eighteen did not and
are honestly absent.
COMPOSABLE BY CONSTRUCTION — the architectural point. Park emits a multiplier
around 1.0, never an additive nudge, because weather has to modulate it next
order: effective = park_base x weather_mod. Additive terms do not compose
correctly (a 5% park and an 8% wind are 1.05 x 1.08, not +13%), and the
challenger converts the multiplier to log-odds so stacking stays correct. A
test multiplies a placeholder weather term onto the park base to prove the shape
composes with no rearchitecting.
DIRECTIONAL BY PROP-OWNER: home_runs and home_runs_allowed both key off hr_base
in the same direction, because the sign lives in the STAT, not the park. Coors
inflates the hitter's home run prop and the pitcher's home-runs-allowed prop
identically.
THREE HONEST STATES, deliberately distinct. Absent (thin sample, adjust
nothing), present (adjust), and weather_na for domes — where the park factor
STILL APPLIES because a dome has a real run environment, and the flag exists so
next order's weather modulation correctly does nothing there. N/A is not absent;
conflating them would either drop a valid park factor or apply wind indoors.
Structural breaks: a season deviating past the threshold starts a new regime
only if the FOLLOWING season confirms it — one odd year is noise, two
consecutive years on the same side is a rebuilt park. Only post-break seasons
are used, so a humidor or moved wall cannot be diluted by the stadium that
preceded it. Factors regress toward neutral by sample size, so a two-season park
cannot assert a Coors-sized coefficient, and fine conditioning stays unavailable
until its own larger floor.
Tests 3669 passed / 297 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
PHASE 0 GATE — historical out-of-sample testing is NOT available, and the reason
matters. statcast_aggregates is overwritten nightly by design (Layer 1 is a full
re-pull upsert), so it holds season-TO-DATE numbers with no point-in-time
history. Classifying a player for a 15 July game using today's aggregate would
feed the model games from 15-21 July — look-ahead leakage, and the resulting
"out-of-sample" verdict would be worthless. The harness therefore reads the
archetype vector RETAINED at grade time (Session 70's instrument) and runs
FORWARD-ACCRUAL, not historical. Reported rather than worked around.
CANONICAL NAMES ASSERTED. Every mapping references the axis keys the classifier
actually emits, and a test walks both maps against BATTER_AXES / PITCHER_AXES.
A key that does not exist would look active and never fire — a mapping that
appears wired while silently doing nothing is the exact failure this guards.
TIER 1 IS LIVE, tautological and directional: PUNCHOUT/WHIFF raises strikeouts;
SINKER/SEAM lowers home runs allowed and FLY BALL/ELEVATOR raises them (a ball
on the ground cannot leave the park); SURGEON ARM/PINPOINT lowers walks allowed;
SLUGGER/BOMBER raises total bases and home runs; TECHNICIAN/SURGEON raises hits
and lowers strikeouts; GRINDER/SNIPER raises walks. Each adjusts only its named
stat, mirrors exactly on the under side, and leaves an average player untouched.
SPEED IS HONESTLY ABSENT. BURNER/stolen-bases has no axis to key on — SB is a
statsapi field that never reached the aggregate store, so Layer 2 shelved it.
The mapping is an empty object rather than an invented one.
THE TIER-2 HARNESS tests MARGINAL CONTRIBUTION, not correlation. A ground-ball
arm obviously correlates with fewer home runs; the question is whether the
archetype explains the PROJECTION'S RESIDUAL (outcome minus p_win). If the
projection already knows it, the residual carries no signal and the mapping is
rejected as redundant — that hurdle is what catches double-counting. The split
is by DATE, never random, because rows from one game share a pitcher, a park and
a lineup and would leak across a random split. Direction is validated from the
held-out data and a contradicted sign is REJECTED, never silently flipped to
whatever the data says, which would be fitting noise.
LIFECYCLE ENCODED — nominated, live, claimed. A mapping that survives runs live
and is measured; only the quantified public claim waits for the ledger. Nothing
sits dark.
One fixture bug worth recording: my first synthetic generator aliased the
carrier selector against the outcome draw and manufactured a 0.038 effect where
the generator had put zero. The harness rejected it correctly — it just gave the
sign reason instead of the redundancy reason, which is how I found it. The draw
now uses a coprime modulus.
Real candidate run end to end, GROUND-BALL to hits-allowed: INSUFFICIENT, 0 of
200 settled rows, because no settled row carries p_win yet (Session 70's
instrument starts recording at the next new lock). That is the correct verdict
and the expected one.
Tests 3654 passed / 296 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The champion (probabilityEstimator -> p_win) keeps serving and grading users,
completely unchanged. The challenger is a second probability computed from the
same inputs at the same instant, landing on the same ledger row so it joins to
the same outcome and the same close. Identical conditions, one difference —
the only clean A/B.
NOTHING IS CLAIMED. Running a challenger is honest beta; asserting it is better
before the settled ledger says so is not. Promotion stays a later decision gated
on Brier and calibration over sufficient segmented volume.
INTERPRETABLE, NOT A RE-ESTIMATION. The challenger is the champion's probability
adjusted by the Layer-2 axes, applied in log-odds space so a nudge cannot push
past 0 or 1 and means the same thing at p=0.5 as at p=0.9. Every deviation is
attributable to a named axis and a signed nudge, stored as
challenger_adjustments, and the total is capped at 0.45 log-odds — a lean on a
real signal, never a re-forecast. Only mechanically obvious stat/axis
relationships are mapped; a speculative mapping would be the same guessing this
layer exists to replace.
IDENTICAL WHERE THERE IS NO SIGNAL, by construction. An unremarkable player, a
thin sample, an unmapped stat or a missing classification all return the
champion's probability byte-for-byte with an empty adjustment list and a stated
reason. The experiment therefore differs only where archetype-awareness could
possibly help or hurt, with no dilution from rows the treatment never touched.
Induced on real players. Judge home runs over: 0.42 -> 0.447, via BOMBER +0.22
and WHIFF RISK -0.11 — two real opposing signals netting positive. The same prop
under mirrors it exactly to -0.027. Judge strikeouts: delta exactly 0, because
WHIFF RISK and GRINDER cancel — an honest "no lean" with both signals still
recorded. Skubal strikeouts over: 0.60 -> 0.702 via WHIFF, TRAPDOOR and CANNON
all aligned; his hits-allowed goes the other way, 0.50 -> 0.392, because a
strikeout arm makes hits less likely. Josh Bell and a 12-PA sample are
untouched.
Isolation is structural: adjust() is pure, the champion field is read and never
written, the served snapshot payload is still the untouched champion object, and
a challenger failure is caught so it can never break the pipeline it is measured
inside. Statcast aggregates load once per snapshot run rather than per prop, so
grade-time I/O stays at zero.
Migration 034. Tests 3634 passed / 295 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Caught by inducing on real rows. The first attach ran and marked 642 rows
market-unavailable while attaching ZERO closes — because it selected a
`fair_prob` column from closing_captures, which has none. That table stores
over_odds and under_odds deliberately (Session 64) so the de-vig can run later
against the same engine the grade-time fair price uses; asking it for a
probability returns nothing and makes every row look closeless.
The de-vig now runs here, via devigTwoWay, which is what makes lock and close
comparable at all. A one-sided capture yields no fair probability and is
correctly not a close.
Repair checked rather than assumed: the 642 markings turn out to be CORRECT —
every one is a game from before closing capture existed on 2026-07-20, so those
rows genuinely have no close and the absence is true. Zero capture-era rows were
wrongly marked. The bug would have mis-marked every future row, which is what
the fix prevents.
Two tests added: the de-vig path with real prices, and a source assertion that
the query never again asks closing_captures for a column it does not have.
Tests 3616 passed / 294 suites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Step 0 found we have been flying without one. p_win lives only in
model_snapshots, which has 1,000 rows and ZERO settled outcomes; the closing
line lives only in closing_captures, which carries no link to a result; and
ledger_entries, the row that actually settles, carries no probability at all.
So "is the projection calibrated" and "does it beat the market" have never been
answerable — the entire measurable universe was 35 rows recovered by a lossy
in-memory join.
PHASE 0 — closing coverage verified BEFORE reuse, because an instrument built
on a partial close measures a biased subset. closing_captures holds 70,254 rows
of which 13,364 are usable, and the 56,890 refusals are candidates we never
graded plus one-sided prices — not refusals of our props. Coverage on graded
props since capture started is 83/83, 100%. Safe to reuse, with the honest
caveat that capture only began 2026-07-20.
THE FOUR-TUPLE NOW LANDS ON ONE ROW. ledger_entries gains p_win, fair_prob_lock,
archetype_vector and projection_locked_at at LOCK time, and closing_prob plus
closing_captured_at from the append-only capture store. The join is the whole
point: calibration is p_win against outcome, market-comparison is p_win against
the close, and both become plain SQL on one record instead of a join that
silently drops 90% of the rows.
p_win and the archetype vector are IMMUTABLE — written once at lock via the
existing ignoreDuplicates upsert, never re-derived at settle. A re-derivation
would measure a projection we never made.
The archetype is stored as the VECTOR, not the label. "Did archetype-awareness
help?" can only be answered against the axes that were live at grade time, and
a single text column cannot express a blend. A grade with no archetype stores
null rather than an empty object.
HONEST-ABSENT BOTH WAYS. A past game with no usable capture is marked
market_unavailable_reason and never given an imputed line; calibration still
scores on those rows, only market-comparison is absent. And a game that has not
started yet is NOT declared closeless — a close can still arrive, and premature
absence is as dishonest as imputation in the other direction.
One bug caught before it shipped: the scheduler hook iterated a SPORTS
identifier that does not exist in that scope. Inside its try/catch it would have
thrown ReferenceError every tick and silently never run — the instrument would
have looked wired and captured nothing. Now iterates cadence.ALL_SPORTS.
The baseline accrues FORWARD. Historical p_win and closes are gone, discarded
before this existed. Calibration and market-comparison stay honest-absent until
volume accrues.
Tests 3614 passed / 294 suites, web build exit 0. Migration 033 applied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
A player is a blend across independent axes, not one label. Skubal is a STARTER
and a strikeout arm and a ground-ball arm and a control arm — four true things
at once, and single-label classification threw three of them away.
AXIS INDEPENDENCE WAS MEASURED, NOT ASSUMED. Correlations over the live store
(467 batters, 531 pitchers); anything |r| >= 0.70 is one underlying trait and
was collapsed so we never show one trait as two archetypes. Batter k% ~ whiff%
+0.89, hard-hit% ~ exit velo +0.88, chase% ~ swing% +0.87, chase% ~ bb% -0.72;
pitcher k% ~ whiff% +0.76, gb% ~ fb% -0.73 — all collapsed.
The survivors are genuinely orthogonal, and one result is worth stating: pitcher
velocity correlates +0.14 with K%, +0.07 with whiff% and +0.07 with GB%.
Velocity is NOT a proxy for missing bats — a hard thrower who misses no bats is
a real distinct type, so CANNON earns its own axis rather than being folded into
STRIKEOUT. Pitcher K% ~ GB% is -0.10, so PUNCHOUT and SINKER are independent,
which is exactly the multi-axis thesis.
Cut-lines are the measured p75 (distinctive) and p90 (elite), per role where the
tails differ even when the medians agree: reliever GB% p90 is 54.1 against a
starter's 48.9, both with a median of 42.5.
THE FALLBACK IS DELETED. classify() used to return FLEX (mlb) / SHIELD (wnba) /
CONNECTOR (nba) at weight 1.0 when nothing scored — "could not classify"
rendered as a fully-confident classification of a real archetype, with
descriptive education copy attached. 8 of 18 MLB players carried it, and FLEX
could never be earned because its only scoring input had zero writers. Every
sport now does what MMA already did: unclassified is absent.
Induced on real players. Skubal: STARTER, throws L, WHIFF + SEAM + PINPOINT, all
elite. Judge: BOMBER + GRINDER + WHIFF RISK — elite power, patient, strikes out,
three true things. Kwan: SURGEON + SNIPER + SLASH with NO power claimed (0.4
barrel% is absent, not "low power"). Josh Bell, who used to classify as DRIVER:
empty blend, "No standout profile — league-average across every measured axis."
Alan Roden, who was FLEX at weight 1.0 on 21 PA: every axis absent, "Not enough
plate appearances yet — no profile claimed."
Per-axis honest-absence holds: a velo-less pitcher keeps every other axis, and
NO DATA is distinguishable from LEAGUE-AVERAGE rather than collapsing into one
shrug. The full vector is stored for Layer 3; only the top three distinctive
traits surface.
Three existing tests asserted the fallback and were updated to assert absence.
One of them surfaced a real robustness gap: classify(sport, null) threw, because
an explicit null does not trigger a default parameter and every scorer
dereferences its argument. Guarded.
Every baseball name is accounted for in docs/ARCHETYPE-AXES.md — built, alias,
tier, or shelved with its unlock condition. Zero orphans; cross-sport names left
for their sport.
Tests 3601 passed / 293 suites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Three Layer-2 prerequisites, each one free call.
HANDEDNESS — statsapi /sports/1/players carries batSide, pitchHand AND
primaryPosition for every player: 1,316/1,316 in the live probe. Batter
handedness was 100% absent, which made every platoon or switch-flavour
archetype unbuildable; it is now populated from the same call that gives
pitchers theirs, with statsapi as the authority and the movement feed as the
fallback.
ROLE — statsapi season pitching with playerPool=ALL returns 751 rows (the
default returns only the ~57 qualified). Real usage: gamesStarted, gamesPitched,
gamesFinished, saves, holds. roleDetail derives starter/closer/setup/reliever
from that instead of the season-IP proxy, which drifts all year as innings
accumulate and left 32 pitchers in a 60-80 IP trough.
VELO — recovered from 53% to 99% (721/729 pitchers). The movement feed carries
only each pitcher's PRIMARY pitch, so matching by position could never do
better than one pitch each. The wide pitch-arsenals feed has one column per
pitch type, matched BY TYPE: Skubal now has velo on all 5 of his pitches. Velo
archetypes are therefore buildable rather than shelved.
Migration 032 applied.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Caught by spot-checking a real row after the backfill landed: Skubal stored
with one pitch. The pitch-movement endpoint with an empty pitch_type returns
ONE row per pitcher — their primary offering — so 677 rows for ~700 pitchers,
and a five-pitch arsenal was being recorded as a one-pitch one. Not a
fabrication, but a silent under-representation of the single most important
pitcher-mechanism field, which is worse than useless for Layer 2: it would have
classified every pitcher as a one-pitch arm.
Mix now comes from pitch-arsenal-stats (3,205 rows = pitcher x pitch type)
carrying usage%, whiff%, K%, put-away% and run value per 100 for every pitch.
Movement still supplies velo, break and handedness, folded onto the primary
pitch; a pitcher present only in the movement feed keeps his handedness and his
one measured pitch rather than being dropped. Velo on non-primary pitches is
null — absent, not guessed.
Skubal now stores 5 pitches, throws L, FF first by usage with velo 96.7.
Tests 3583 passed / 292 suites.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Found by inducing the real job on the server, not by review: the first chunk
wrote, the second failed with 'ON CONFLICT DO UPDATE command cannot affect row
a second time'. A player can legitimately appear in BOTH the batter and the
pitcher feeds — two-way players, position players who pitch, pitchers who bat —
so (sport, season, source_id) collapsed two real profiles into one key and a
single batch hit the same row twice.
Ohtani has a real batter profile and a real pitcher profile. Merging them would
invent one player out of two genuinely different sets of measurements, so role
goes in the primary key rather than one profile winning. Migration 031 applied;
conflict target updated; a two-way case is now a test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The data foundation for the archetype and projection layers, built as the
pattern every sport inherits. Layers 2 and 3 are not touched.
PHASE 0 GATE — both match rates measured live, both 100%. Batters 40/40;
PITCHERS 66/66 across five real rosters (CLE, DET, MIN, NYY, LAD) joined by
MLBAM id against the 713-pitcher Savant feed. Zero honest-absent on identity,
because the join is an integer both systems use natively — and the snapshot
pipeline already stores it per graded row.
SOURCE — five Baseball Savant CSV leaderboards, free and public, pulled with
axios and the CSV parser savantAdapter already runs in prod. pybaseball is
deliberately NOT used: it is an MIT wrapper over these same URLs, and adding it
would reintroduce a Python runtime in a stack where the existing Python service
is already offline. min=1 on every feed, not Savant's default min=q, so the
long tail arrives and OUR minimum-sample gate decides what is thin — explicit
and testable rather than silently dropped upstream.
Measured: 1,354 rows per season (604 batters, 750 pitchers), all five feeds in
about five seconds. Pitcher mechanism includes arm angle, GB/FB/LD, chase and
whiff; batters get exit velo, launch angle, barrel and hard-hit, chase and
z-swing. Handedness rides in free on the movement feed (677 pitchers); batter
handedness stays absent pending a roster join rather than being guessed.
BACKFILL AND REFRESH ARE THE SAME CALL — a full re-pull upserted on
(sport, season, source_id). Idempotent and self-healing: a missed night
self-corrects on the next run, with no incremental who-played bookkeeping to
drift out of sync. At 1,354 rows the simple thing is also the robust one.
HONESTY RULES, each with a test: a metric the feed did not carry is null and
never 0; a thin sample is STORED and flagged rather than dropped or inflated,
because thin and missing are different claims; an unjoined player is stored
with a null player_key and joins later; and if every feed comes back empty the
job REFUSES to write, so a bad night can never blank a good table.
Freshness is treated as a truth property. updated_at on every row, and the
scheduler pages on a failed run AND on silent staleness — a job that stops
being scheduled never produces a failure, so staleness has to alarm on its own.
Never-built is deliberately not stale: different condition, different fix, and
paging on a fresh install teaches the operator to ignore the alarm.
Nightly at STATCAST_HOUR_UTC (default 11 UTC, after every game is final), kill
switch STATCAST=0, and induce-able at POST /api/internal/statcast/refresh with
a freshness probe at /statcast/status — we verify a refresh by running it, not
by waiting for the slot.
Migration 030 applied. Promoted columns for the classification-critical metrics
plus a metrics JSONB carrying every raw field, so Layer 2 can reach something we
did not promote without a re-ingest. Raw per-pitch stays out of Postgres on
purpose: one season is ~0.85 GB against a 500 MB plan ceiling, and it is
re-pullable from the free source if Layer 3 ever needs it.
Pattern documented in docs/MECHANISM-DATA.md for NBA tracking and NFL Next Gen.
Tests 3581 passed / 292 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Live induction on the deployed landing page caught this; markup review and the
unit suite both passed it. With the model leg stripped for anonymous visitors,
LiveHeroProp forwarded book/fair/model/ev/quarantine to PriceTriplet but NOT
model_price_locked, so deriveValueState fell through to the missing-model-price
branch and the card rendered "MODEL READ WITHHELD" — the quarantine state,
whose copy says we suppressed our own price because a leg is poisoned. Nothing
was poisoned. The real reason was the paywall, and the two must never share a
face: one says our data is untrustworthy, the other says you don't have this
tier.
Now forwarded, with a test asserting both the separation in deriveValueState
and the forwarding at the call site.
Tests 3559 passed / 291 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
The landing hero reads the snapshot from Redis DIRECTLY via heroPropService,
so it never passed through routes/snapshot.js and was still serving model_odds,
ev_pct, value and takeable to anonymous visitors after the first fix. Same
strip, same tier resolution, same private-cache rule for authenticated callers;
the Next proxy now forwards the bearer token.
PRODUCT CONSEQUENCE, FLAGGED RATHER THAN BURIED: the landing hero is served to
anonymous visitors, so it now renders BOOK and FAIR with the model leg LOCKED
instead of the full triplet it showed this morning. That follows the stated
free-tier rule exactly, but it does trade a strong marketing moment (VALUE
+21.1% VS FAIR on the shop window) for consistency of the gate. Reversing is
one line — add 'model_price' to the free tier in src/config/tiers.js, or
special-case the hero route — and is a product call, not a correctness one.
Tests 3557 passed / 291 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
PHASE 0.5 GATE — the three checks, and one correction.
`fairLine` does not exist. Zero hits across src/ and web/src. Option A as
written had no referent, but it resolves better than feared: `fair_odds` is
already a real de-vigged American price on every graded snapshot row, so there
is nothing to derive.
Gate 1 (is it a price): PASS. fair_odds is American odds from
impliedProbToAmerican inside devigTwoWay; fair_prob is the probability. Both
distinct from `line`, the stat threshold.
Gate 2 (numeric match): PASS, 8/8 exact. Recomputed fair_odds and fair_prob
independently from the stored raw over/under prices; every value matched the
stored one to the integer and to 3dp. Same de-vig, same numbers the component
was proven against.
Gate 3 (poison independence): PASS, and proven on the quarantined cohort
itself. devigTwoWay's inputs are (over_odds, under_odds) — market prices
only, no model term is reachable. The 8 rows recomputed above are all
wrong_opponent_grade rows, and their fair prices reproduce exactly from the
market. The poison is in the grade, not the price. Quarantine therefore
suppresses the MODEL leg only; the fair leg stands, as designed.
THE LEAK WAS REAL AND ALREADY LIVE. GET /api/snapshot/:sport is public and
unauthenticated, and it was serving model_odds, p_win, ev_pct, value and
takeable to anonymous callers on every graded row — 25 of 25 on the live wnba
board. The Session-66 gate on /api/analyze was bypassed entirely by this
endpoint.
The strip covers more than model_odds, because model_odds is not the only way
to read the model price: p_win IS the price in another base, and ev_pct is
INVERTIBLE — ev is a function of p_win and book_odds, and book_odds is public,
so leaving ev behind hands the price over. All five model-derived fields go.
book_odds, fair_odds, fair_prob, overround and devig_method stay on every tier:
the fair leg is never the paywall. Rows that keep a book+fair pair are stamped
model_price_locked so a gated price is never mistaken for a missing one.
Tier comes from resolveTierFromRequest, which reads a bearer token when one is
present and otherwise returns 'free'. It FAILS CLOSED on every error path, so a
resolution failure can only ever withhold the price. The response now varies by
entitlement, so the /:sport handler downgrades Cache-Control to private for
authenticated callers and the browser proxy forwards the bearer token —
otherwise a CDN could hand a paid payload to an anonymous viewer, or every
request would look anonymous and paid users would lose the leg.
READ CARD — a manual scan carries no market. The request is {player, stat,
line, direction}, so the engine has no over/under prices to de-vig and
book_odds/fair_odds are legitimately absent from its response; that is why the
triplet was hidden there. lookupSnapshotPrices recovers them from the
pre-graded snapshot via the same cache-only read this route already performs
for locked odds and team. The join is exact on player + stat + line + side
(fair_odds is side-specific), and returns nothing unless book and fair are BOTH
present — a user-chosen line the board never graded has no market attached, so
the triplet stays hidden rather than borrowing another line's price.
FAIR-LEG ABSENCE, measured before shipping: 636 graded rows, 636 with book,
636 with fair, 0 one-sided. Absence rate 0.0%. The hero number is not a
sometimes-number on current data.
Tests 3556 passed / 291 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
PHASE 0 finding, reported before building: the token layer this order asked
me to establish ALREADY EXISTS and already matches HANDOFF.md exactly.
web/src/app/globals.css :root carries the design's surfaces, borders, text
ramp, fonts and grade colours byte-for-byte (aligned 2026-07-16), and
lib/colorContract.js already encodes the green-is-edge-only and
glow-is-A-tier-only laws with a test enforcing them. The stack is Tailwind v4
CSS-first (no config file) with components styled by inline style={{}} reading
var(--x) — 1,916 such reads — so CSS custom properties are the only vehicle
the stack natively consumes. Creating a second parallel layer would have meant
two competing sources of truth, so this EXTENDS the existing one.
PHASE 1 — additive only. globals.css gains one colour the system did not have,
the priced-out blue (#8fb2de + tints), plus a tokenized A-tier glow and the
fair-leg tints. The block writes the LAWS into the token layer itself — green
= takeable edge only, glow = A-tier only, amber = caution + the fair leg, red
= miss/negative only, blue = edge priced out, JetBrains Mono = all data — and
a test asserts every newly-declared name is new (zero collisions, zero
overrides). No existing hardcoded style was touched and no live surface was
migrated: the diff over existing files is 149 insertions, 0 deletions.
PHASE 2 — lib/valueState.js is the single verdict function; the component
renders what it returns and never re-derives one. VALUE fires only on
ev >= 2 AND a takeable price, mirroring src/config/valueEngine.js with a test
that cross-checks both files and fails on drift. Five states: VALUE (green),
EDGE-NOT-TAKEABLE (blue), NO EDGE (grey, stated at full voice), QUARANTINE
(model leg withheld, book+fair stand), REFUSAL (nothing rendered). Free tier
is gated at the wire — tierGating strips model_odds and sets
model_price_locked, so the lock is real rather than a blur over data already
sent; book and fair pass through on every tier because the fair leg is never
the paywall. Wired into the landing hero (data was already on /api/hero-prop)
and the read card, where the projection block reads first and the triplet sits
beside it, not in place of it. Ledger and public profile are out of scope —
no fair-odds columns exist there.
PHASE 3 — induced all six states in a real browser and read back computed
styles, not just markup. Green resolves on VALUE alone: rgb(0,212,160) on the
model leg and verdict; the +11.7%-EV-at-210 row renders rgb(143,178,222) blue
and a white model leg; quarantine shows MODEL "—" with book and fair intact;
refusal renders no legs at all; free tier renders a lock bar with book -120 and
fair -104 still honest. Landing hero on live data: book -153, fair -129, model
-343, VALUE +21.1% vs fair at +28.1% EV. At a real 390px column the three legs
hold at 117px each with no horizontal overflow and fair no smaller than its
neighbours.
Induction caught a real bug that markup review would not have: the "VS FAIR"
figure compared BOOK to fair, printing "VALUE · -6.5% VS FAIR" — a
contradiction on screen. The design's own two worked examples pin the formula
as MODEL minus FAIR in implied-probability percentage points; modelVsFair now
reproduces both exactly (+2.9 and -1.8) and a test locks them. The figure is
shown only when its sign agrees with the verdict, so a row that clears the EV
bar on the book price while our price sits level with fair leads with the EV
instead of a number that reads as a contradiction.
Tests 3539 passed / 290 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Drops the updated claude.ai/design bundle into specs/design-reference/ as the
build reference. New since the Jul-18 export: Vyndr Price Triplet.dc.html (the
reality-corrected ACT 01 — five honesty states incl. EDGE-NOT-TAKEABLE, and the
projection-then-price read-card hierarchy), Vyndr Offseason.dc.html,
Vyndr Intelligence.dc.html, PNG export masters, and 83 archetype glyphs
(74 display + 9 classifier-legacy).
Reference material only. No product code touched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
Two Truth-Law fixes found by auditing the product logged-out.
FIX 1 — /u/[handle] claimed a "CLV-verified record" with "closing-line
value included" while ZERO closing-line value renders there. Verified
live: GET /api/profiles/vyndr returns beat_close_pct null (gated behind
CLV_CAPTURE_RELIABLE, unset while C4 is open). Eight instances found —
two of them (the OG + portrait "CLV-VERIFIED RECORD · 30D" eyebrows)
only by the post-removal residual sweep; two more printed the claim in
exactly the no-record branch.
Copy now describes what the page shows. The gated CLV-VERIFIED badge and
the BEAT CLOSE figure are removed from the public profile, OG card and
portrait card. DISPLAY ONLY: beat_close_pct, clvCaptureReliable() and
the whole CLV data path are untouched, and the earned directional badge
stays Analyst+Desk. The claim returns when CLV genuinely renders here.
Also fixes the doubled "· VYNDR · VYNDR" title (layout's '%s · VYNDR'
template already supplies the suffix); verified on composed output by
serving the build and reading the real HTML, not on source.
FIX 2 — the player page's FORM was `70 + 4 × (count of tonight's graded
props)`. Nothing on the HTTP path ever sets stats.form, so that fallback
WAS the live number: Josh Bell's "74" is 70 + 4×1 prop, confirmed
against his live payload. MATCHUP was gradeFromForm(that number), with a
hardcoded 'B' on the no-archetype branch — both fabricated letters with
no opponent input on the path. Systemic: buildIntel is the unconditional
path for every player and sport.
FORM and MATCHUP now render "—" (kind 'plain', so no bar width or colour
is computed off a null). gradeFromForm is deleted and the prop count is
no longer passed into buildIntel. computeFormScore's hardcoded 75 now
returns undefined. Induced across MLB/NBA/WNBA: all render cleanly, and
real values (USAGE 3.6 AB/G, REST B2B) still render.
Neither form value feeds the grade — engine1 reads raw l5_avg/l20_avg
against the line and never a form key; buildIntelFields decorates the
already-graded object. Grade inputs are byte-identical.
Held (needs a per-sport headline-stat design call): a real player-level
form metric + label disambiguation.
Tests 3491 passed / 289 suites, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj
TRUTH FIX FIRST. row.clv/clv_result derive from the overwritable
closing_line — 637/699 rows had closing == locked — so of the 578 rendered
chips, 522 "flat"s encoded a CAPTURE FAILURE as a held line. That number
was live on FIVE surfaces, not the one the order assumed:
1. ledger card ClvChip@319 (authed)
2. public profile /u/ duplicate chip@260 (PUBLIC, shareable)
3. dashboard "· CLV BEAT"@508 (authed)
4. board / Slate "CLV BEAT/FADED"@422 (FREE surface)
5. board / Slate "✓ HIT · CLV BEAT"@517 (FREE surface)
Removing it from the ledger alone would have left the lie live on three
surfaces including two public ones, defeating the stated PRIMARY GOAL, so
the removal covers all five. That is a deliberate extension beyond the
"scoped to ledger card" guardrail and is flagged as such — the guardrail
protected against feature creep, and this is the same defect at four more
addresses.
DIRECTIONAL BADGE installed in the old ledger slot. Three time states now
read distinctly on the row: entry (locked_odds, at-grade) · close (badge,
at-close) · outcome (hit/miss, final). The RESULT stays the row hero.
The PUBLIC profile deliberately gets NO badge — it never receives dclv
data (Analyst+Desk, server-gated), so that surface now shows the settled
result alone.
HONEST ABSENCE, tested: the 578 formerly-chipped rows now render NOTHING —
not a "flat", which is the old chip's lie in subtler form. Rows settled
before capture existed will never get dclv, and permanent silence is the
correct output. All six states verified in place; a 40-row dense ledger
stays legible (27 badges, max 40 chars, one line, consistent slot after
OutcomeChip); no CLV sort or filter exists.
FIELD REMOVAL DEFERRED as a separate scoped cleanup: /api/ledger/mine and
/api/profiles still SEND clv/clv_result, and dashboard/Slate still type
them. Display removal is local and safe; stripping the fields mid-swap
could break a response consumer.
Suite 289/3485 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
CARD BADGE ONLY. Ticker-CLV explicitly DEFERRED (named, not lost).
PHASE 1 — SERVER-SIDE GATE AT THE DATA LAYER. A Free request never
RECEIVES dclv data: the CLV columns are appended to the SELECT only behind
canAccess(tier,'clv_badge') (new capability, analyst+desk), and responses
are ALSO stripped as defence in depth so a future SELECT change cannot
quietly leak. No CSS/client gate — data that reaches the browser has left
the building. dclv_fair_lock/fair_close are de-vig internals and are never
sent at all.
SURFACE AUDIT, all six channels, each test-locked to contain no CLV:
public profile (share link), snapshot/card feed, ticker feed, share
card/OG, embeddable widget, newsletter. A test also asserts no
ledger_entries read uses select('*') — a star would auto-leak every new
column, which is exactly how a gate becomes theatre.
PHASE 2 — IMMUTABLE ONCE COMPUTED. A settle can re-run (stat correction,
protested game) and a badge that flips positive->negative AFTER a user saw
or screenshotted it is a credibility failure. First computation wins: dclv
is only computed when dclv_computed_at is null, so a re-settle can never
rewrite a shown badge. Same discipline as the locked grade.
PHASE 3 — RENDER, test-first, ABSENCE IS HONEST. unknown / flat / null /
missing-receipt all render NOTHING — no element, no placeholder, no
"pending". Proven on an ALL-NULL board (today: 0 badges) and a MIXED board
(tomorrow: 1 of 4 badged, badge-less cards clean). Binary states only:
positive -> MOVED TOWARD US "graded -110 · closed -145"
negative -> MOVED AWAY "graded -110 · closed +120"
The RECEIPT is the persuasive part, so a badge with no numbers is
suppressed rather than shown as a bare claim. Negative is neutral context
and NEVER touches the locked grade — no back-door re-grading.
NO aggregate, count or rollup exists by construction: the module exports
exactly {clvBadge, fmtPrice} and a badge payload carries exactly
{tone,label,receipt} — asserted by test, because an on-screen tally would
be the held aggregate claim through the side door.
Build gotcha hit and fixed: clvBadge is CommonJS (allowJs) with no TS
types, so the .tsx needed an explicit cast at the call site — the build
worker exits 1 on type errors even though compilation "succeeds".
Suite 288/3473 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
PER-READ ONLY. No aggregate CLV stat, no CLV marketing un-held.
PHASE 0 FINDING THAT SHAPED THE BUILD: the LOCK end must come from
model_snapshots, NOT ledger_entries. The ledger stores only the graded
side's locked_odds (694 rows, single-side) which CANNOT be de-vigged.
model_snapshots retains BOTH side prices on 520/520 graded rows AND an
already-de-vigged fair_prob on 520/520 — produced by the same
devig.devigTwoWay the close uses, so "same method both ends" holds by
construction rather than by convention.
THE COMPUTE TRIGGER is the SETTLE PASS (ledgerService.settleLedger). At
settle the game is final, so the close has landed and the read is final —
the only moment both ends of the comparison exist. Grade and locked prices
are written hours earlier and the close at lock, so without this trigger a
correct CLV function would simply never populate.
JOIN INHERITS THE PROVEN KEY: (sport, player_key, stat, side, game_date),
WITHOUT line — a close that moved off the graded line is the entire point.
Verified clean earlier: 164 identity groups, zero ambiguity. Rows whose
capture refused (missed/ambiguous/one-sided) are UNKNOWN for CLV, matching
the capture layer's own honesty.
SIGN IS SIDE-BOUND and proven by test before the logic existed — the
badge-inverting trap. Same market move:
OVER-graded -> positive clv +0.0800 (fair .500 -> .580)
UNDER-graded -> negative clv -0.0800 (fair .500 -> .420)
exact mirrors. FLAT is a PROBABILITY-space threshold always (1.5pp): a
40-cent price move on a deep favourite reads flat, correctly, because
price space lies about magnitude.
UNKNOWN is a first-class state, never 0 — zero asserts "the market did not
move", which is a claim; a missing close asserts nothing. describe()
returns null for unknown so a badge can never render for it.
migration 030 adds dclv/dclv_state/dclv_fair_lock/dclv_fair_close/
dclv_computed_at as NEW columns rather than reusing the C4 clv fields —
conflating a verified per-read signal with a known-broken one would be the
worst kind of quiet lie.
Caught pre-deploy: the trigger call passed `deps`, which is not in scope in
settleLedger (it uses `opts`) — a ReferenceError at the call site, outside
the helper's try/catch, which broke two settlement suites.
Suite 286/3447 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
featureCache.teamFeatures now derives MLB opp_rank_stat from statsapi team
pitching splits when the ESPN path yields nothing — which for MLB is
always, because ESPN's MLB team endpoint carries no defensive metric at
all. mlbStatsAdapter.getTeamPitchingStats fetches all 30 teams in one free
unauthenticated call, cached at the season TTL.
CONSUMPTION PATH VERIFIED before wiring, not assumed:
featureCache.teamFeatures sets out.opp_rank_stat (line 338)
-> engine1.computeFactors READS features.opp_rank_stat (lines 96-102)
-> fires weak_opponent_defense (>=0.70) / top_opponent_defense (<=0.30)
So teamFeatures is the correct insertion point: the grader reads exactly
the field we populate. A value written anywhere else would have been a
dead end — computed, retained, and still not affecting the grade.
Contract preserved: the derived value goes into the SAME field with the
SAME 0-1 scale and the SAME high=weak polarity WNBA uses, so engine1 reads
one field with one meaning across sports. Isolated and best-effort — a
derivation failure leaves the field ABSENT (honest null), never a guessed
rank. Only fills when the ESPN path produced nothing, so WNBA behaviour is
untouched.
Suite 285/3435 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
PHASE 1 — CLOSE-CAPTURE RETRY, test-first. The closing capture gets a
retry the snapshot path deliberately does not: a snapshot re-runs at the
next slot, but a MISSED CLOSE IS PERMANENT, and the feed flaked once on a
dry induce. Three hard rules, each driven by a test written before the
logic:
- BOUNDED attempts (default 3) with short backoff so every attempt fits
inside the window. Never infinite.
- HARD LOCK-WALL: inside lockWallMinutes of first pitch (or past it) it
stops and records missed_close. A price captured AT or AFTER lock is
NOT a close; storing one would fabricate the CLV baseline.
- NO BOUND LOCK TIME -> refuse immediately, never burn retries on a prop
whose close cannot be timed.
On exhaustion it records missed_close with NO price — never a stale,
mid-day or post-lock line.
PHASE 3 — MLB opp_rank_stat DERIVED, contract-locked. MLB previously had
no opponent metric at all (ESPN's MLB team endpoint carries none), so
engine1's +/-1.0 opponent factor never fired for the sport carrying most
of our volume. Derived from data we already ingest: statsapi team pitching
splits, all 30 teams in ONE free unauthenticated call.
THE SHARED CONTRACT is documented and TESTED, not assumed: 0-1 scale,
HIGH (>=0.70) = WEAK opponent, LOW (<=0.30) = TOUGH — identical to WNBA's
live semantics. Polarity is the highest-risk part: backwards polarity does
not fail loudly, it silently adjusts every MLB grade the wrong way. A test
asserts MLB polarity EQUALS WNBA polarity using engine1's own thresholds.
PROVEN AGAINST THE LIVE FEED:
Colorado Rockies BAA .286 -> opp_rank 0.983 (weak, fires weak_opponent)
LA Dodgers BAA .215 -> opp_rank 0.017 (tough, fires top_opponent)
POLARITY HOLDS: true
HONEST NULLS, tested: thin league baseline, thin opponent sample, unmapped
stat, unknown opponent, or a missing field all return NULL with a reason —
we are FIXING a silent null, so it is never replaced by a confident guess
off three games. opponentStrengthHealth pages on an empty source AND on
derived-null-for-a-sport-we-expect-to-derive.
Suite 285/3435 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
PART A — HARNESS ARMED ON OUR OWN INFRA. snapshotScheduler now runs the
nightly backtest at HARNESS_HOUR_UTC (default 14), appends to
harness_results, and pages via opsWatch.harnessStaleAlarm — a validator
that stops running looks exactly like one that keeps passing. No external
dependency: the join is plain SQL through the service client and the
harness is a pure function. POST /api/internal/harness/run induces the
same code path on demand, because a scheduled mechanism is verified by
inducing it, never by waiting for a slot.
PART B PHASE 0 — GATE PASSED for what is capturable:
- C4 diagnosed: closing_line is ONE overwritable field with no timestamp
and no provenance. captureClosing writes the current line and, when a
prop fails to match, silently leaves the earlier value (= the lock) in
place — so "captured a real close" is indistinguishable from "never
updated". It is 92% equal, not 100%: 56 rows DID record movement, so
the defect is provenance, not the value.
- Feeds: normalized props already carry BOTH raw side prices per book,
with game_time, and the intraday refresh polls every ~20 min during
slate hours — so the last observable pre-lock line is available.
- SHARP close: pinnacle is in ALLOWED_BOOKS -> a no-vig reference is
capturable ("beat the market").
- ODAWA: NOT capturable. 'odawa' exists only as a UI preference option in
onboarding/settings; it is in no adapter, no ALLOWED_BOOKS, no feed. An
un-capturable source is a finding, not a gap to paper over.
- JOIN: must drop `line` from the natural key, because a close that MOVED
off the graded line is the entire point of CLV. Verified safe — all 164
current identity groups have exactly ONE line per
(sport, player_key, stat, side, game_date). Zero ambiguity.
PART B PHASE 1 — CAPTURE ONLY, built test-first. The refusal was proven
before the capture logic existed: unbound game_time, doubleheader
ambiguity, a missed pre-lock window, or a one-sided price all record
missed_reason with NO price. A stale or mid-day line substituted for a
close would manufacture a CLV proof from a number that was never the
close.
migration 029 closing_captures: append-only, never overwritten (that is
the provenance C4 lacked), BOTH raw side prices so the existing de-vig
engine can compute a fair closing probability later, sharp vs book line
types kept distinct. Wired into the intraday refresh with a capture-rate
alarm — a missed close is unrecoverable.
NO CLV metric built, as ordered. This starts the clock.
Suite 284/3417 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Phase 0 gate PASSED: the join is clean. No FK exists; the natural key
(sport, player_key, stat, line, side, game_date) yields 283 clean 1:1
joins with ZERO ambiguity. game_id is NOT usable — 400/550 snapshot rows
carry UNK@UNK because home/away names weren't threaded into the grader
until Order 1.6. Non-joining rows are EXPECTED, not errors: retention
stores both sides plus refusals; the ledger keeps only the graded side.
Outcomes are NOT denormalized — ledger_entries stays the source of truth.
BUILT TEST-FIRST, and the first property proven is the REFUSAL, not the
math. Below threshold the harness emits INSUFFICIENT with n and the
shortfall and NO rate anywhere in the payload, so a downstream renderer
cannot surface one by accident. A test asserts the payload contains no
hit_rate number at all.
- Wilson intervals (correct at the n we actually have, unlike the normal
approximation which emits negative lower bounds).
- Strata NEVER mix sport or model_version.
- Denominator excludes quarantined, void, unrecoverable, pending, push —
asserted by test.
- Monotonicity refuses to RANK buckets whose intervals overlap; it reports
"not distinguishable on this sample".
- Probability calibration (Brier + reliability) also respects the
threshold: a thin sample returns status INSUFFICIENT and a NULL score.
- Replay seam reads the STORED feature vector only. A row whose input was
never retained is UN-BACKTESTABLE, never scored with substituted current
data. Identity replay reproduces the live prediction exactly.
The tests caught a real bug in my own code: `Number(null) === 0` let a
null p_win through as a confident 0% forecast — this codebase's signature
fabrication bug, inside the harness whose entire purpose is refusing
invented numbers. Fixed with a strict null guard.
FIRST LIVE RUN — the correct, passing output:
VERDICT: INSUFFICIENT_HISTORY (can_validate=false)
283 joined -> 35 scored (120 quarantined, 124 pending, 4 terminal)
C n=18 (short by 2), B n=17 (short by 3)
strata: mlb 7, wnba 28 — never mixed
migration 028 adds harness_results (append-only trend log; INSUFFICIENT
rows are expected and correct) and opsWatch.harnessStaleAlarm pages if the
harness stops running — a validator that isn't running looks exactly like
one that keeps passing.
Suite 283/3403 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Order 2 Phases 2 + 4. Pre-heal rollback point secured first:
vyndr-20260720-093821.dump (856,890 bytes) VERIFIED ON THE BOX, not just
exit 0.
MIGRATION 027 — two DISTINCT exclusion scopes, deliberately separate:
- quarantine_reason: the row's GRADE is untrustworthy (wrong_opponent_grade).
The row REMAINS a real public settled result — the bet happened, the
outcome is real — but it must never train or validate, so
getModelAggregate now excludes it from the denominator alongside
void/unrecoverable.
- analysis_flags: the row is VALID for settlement and the record but
unattributable for PER-GAME analysis (doubleheader dates). Explicitly NOT
filtered from aggregates.
Collapsing these would either wrongly drop 166 doubleheader rows from the
record or wrongly keep 25 wrong-opponent grades inside model validation.
Tests assert both directions, including that analysis_flags is NOT filtered.
Also adds re_settled_at + settlement_source to model_snapshots.
DNP VOIDING RE-ENABLED — reversing my own Order 1.5 disable, with scrutiny,
because its premise was FALSE. Order 1.5 assumed a missing player row meant
the row's DATE was wrong. The Phase 0 dry-run disproved it: across every
bindable row the stored date matched a real game (MIS-DATED: 0), and the
players I had cited as counter-evidence were genuine DNPs on their true
dates (Freeman 07-18; Kwan/Hedges/Davis 07-17 — their teams played, they
did not). The evidence is positive: games FINAL + no line in a full-season
log = no bet existed.
I got this wrong twice tonight in opposite directions; the dry-run is what
caught it. Recording the reasoning in the code so the next reader sees why
the flag flipped back.
Suite 282/3386 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Order 1.6 Phase 1. This is a MODEL-OUTPUT fix, not bookkeeping.
computeFeatures.lookupTodayGame called the ESPN scoreboard with NO date
param and took whatever ESPN calls "today". Renamed to lookupGameOnDate
and now sends ?dates=YYYYMMDD from the prop's BOUND game — the same game
the ledger, retention and settlement use, so all four finally agree.
PROVEN against live ESPN (before/after, same instant):
dateless "today" CLE->PIT NYY->LAD LAD->NYY (Jul 19 card)
bound to 2026-07-20 CLE->MIN NYY->PIT LAD->PHI (the real games)
bound to 2026-07-19 CLE->PIT NYY->LAD LAD->NYY (reproduces OLD)
Every opponent was wrong. opponentAbbr feeds opp_rank_stat (a +/-1.0
factor) and isHome feeds home_away (+0.5), so late-slot grades were
scored against the wrong matchup.
Note the window is WIDER than the 01:00/03:00 UTC slots: this ran at
07:5x UTC = 03:5x ET and ESPN's dateless scoreboard was STILL returning
the previous day's card.
HONEST DEGRADATION: with no bound game date the grader does NOT fall back
to a dateless lookup — it records 'no_bound_game_date' and leaves
opponentAbbr/isHome/gameId null, so engine1 simply omits the opponent and
home/away factors rather than scoring a wrong matchup. Tests assert both
directions.
Same class of bug fixed alongside: the Tank01 augmentation used TODAY's
UTC date for its cache key; it now uses the bound game date.
gradeSlateService threads game_date/game_time/home_team/away_team into the
grader so the binding reaches computeFeatures at all.
Audited the rest of the feature path for dateless/"today" lookups — none
remain (weather is current-conditions by venue, park/pace are static).
Suite 282/3383 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Order 1.5 Phase 1. PropLine emits NO commence_time (grep-verified: zero
hits in proplineAdapter), so ledgerService's
`dateET(prop.game_time) || dateET(gradedTs)` always fell through to the
GRADE timestamp — and a 01:00/03:00 UTC snapshot is 21:00/23:00 ET the
PREVIOUS day. Tonight's props were filed under yesterday, settlement
correctly found no game there, and Order 1's void logic turned that into
64 destroyed results.
gameBinder.attachGameTimes() now matches every prop to a scheduled game by
TEAMS across the plausible ET window (grade date, +1, -1) and attaches the
GAME'S OWN time/date/id. It runs in snapshotService before grading and
before the ledger write, so ledger, retention and settlement all inherit
the correct date from one place.
HARD CONTRACT: an unbindable prop returns NOTHING. ledgerService no longer
has a grade-clock fallback — a row with no real game time is SKIPPED and
counted, because a mis-dated row is fabricated data and the ledger holds
real values or nothing. A slate that binds nothing pages.
DOUBLEHEADERS are reported, never guessed: two games with the same teams
on one date mark the binding `ambiguous` so settlement can decline rather
than attribute a prop to the wrong game. (Real example already in the
data: mlb:2026-07-11:MilwaukeeBrewers@PittsburghPirates(Game1).)
Also fixes retention, which had the SAME bug from last night — I had
dated model_snapshots rows with the snapshot clock. Rows now take the ET
date of the bound game_time.
Suite 282/3381 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Correctness fix to code I shipped minutes ago. The induced live settle
pass voided 64 rows as 'player_dnp' and a large share of them are WRONG:
the Jul 18 set is everyday starters (Freeman, Bellinger, Tucker, Chisholm,
Conforto). They played.
ROOT CAUSE — and my Phase 0 diagnosis was wrong. It is not DNP. The
ledger row's game_date is WRONG. ledgerService derives game_date from the
GRADE timestamp when the feed carries no game_time, and a 01:00/03:00 UTC
snapshot is 21:00/23:00 ET the PREVIOUS day, so rows get labelled with the
previous ET date. Verified against fresh season logs (cache disabled, so
not staleness; found:true, so not name resolution):
Freddie Freeman played Jul 17 and Jul 19 (x2, doubleheader) — NOT Jul 18
Steven Kwan played Jul 18 (x2) and Jul 19 — NOT Jul 17
Settlement was correct to find no game on the labelled date. My void logic
then converted a data-labelling bug into destroyed results.
FIX: never void on player-absence alone. Voiding now requires POSITIVE
evidence — the games themselves postponed/cancelled. Absence returns
'unknown' (reason player_absent_unconfirmed), so the row retries and ages
out to 'unrecoverable' at the cap. We cannot distinguish "did not play"
from "mislabelled date", so we must not claim DNP. Both terminal states
are excluded from the record denominator either way.
Window-decay remains genuinely fixed (full season log vs a rolling
window), and terminal states still prevent immortal rows.
NOT DONE HERE: the 64 wrong voids are still in the table, and the
game_date derivation is still wrong at the source. Both are reported for
the table — no healing in this order.
Suite green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Order 1 of 2. Push scoring UNTOUCHED — it is correct. No healing here.
PHASE 1 — DATE-TARGETED FETCH replaces the rolling window for settlement.
settleSource.resolveOutcome() resolves the SPECIFIC DATE and, when the
player is absent, reads GAME STATE to learn what the absence MEANS:
game final + player has a line -> SETTLE (a partial game is a real
result, never a void)
game final + player absent -> VOID (confirmed DNP)
postponed / cancelled -> VOID
scheduled / in progress / SUSPENDED -> PENDING (a suspended game resumes;
voiding it would destroy a real bet)
player played, stat missing -> unknown, NEVER void a real appearance
This is FREE for MLB: mlbStatsAdapter.getPlayerGameLog already returned the
full season log and getPlayerStats was discarding it with .slice(-10).
Settlement now reads fullLog — same request, same cache — which removes
window-decay entirely (the verified failure was a Jul 12 game outside a
last10 starting Jul 6). Projections keep using last10, unchanged.
PHASE 2 — TERMINAL STATES (migration 026 applied). outcome CHECK widened to
hit/miss/push/void/unrecoverable; added settle_attempts, settlement_source,
settlement_version, model_version. A row that cannot be resolved after
SETTLE_ATTEMPT_CAP (4) date-targeted attempts becomes 'unrecoverable'
rather than pending forever. CRITICAL: getModelAggregate now EXCLUDES void
and unrecoverable from the settled selection — it used
.not('outcome','is',null), so without this a void would have counted as a
settled row and silently moved the public record. Verified in the record
calc, not just the settle path.
PHASE 3 — SETTLEMENT-RATE ALARM. zeroSettleAlarm only caught a TOTAL zero
while ~30% of a slate failed quietly (Jul 17: 57/86). opsWatch
.settlementRateAlarm pages when resolved/attempted falls below
SETTLE_RATE_FLOOR (0.8). Voids count as RESOLVED — a void is a legitimate
terminal state — so healthy voiding never pages. Third silent-failure
surface of the night, now closed.
PHASE 4 — VERSION STAMPING. src/config/modelEras.js defines the cutoff
ONCE (2026-07-19T22:50:00Z); migration 026 backfilled pre-cutoff rows as
'pre-retention-unknown' (naming the uncertainty, not implying knowledge);
new rows carry model_version.
Regression caught pre-deploy: getScheduleFn was not injectable, so the
ledger suite hit the real network and HUNG. Now injectable via opts and a
no-op under NODE_ENV=test. The "no row -> pending" test was updated to the
new behaviour deliberately: a missing row on a FINAL game now voids.
Suite 281/3373 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
PHASE 1 — cron capture needed NO wiring. Verified in code: the scheduler
tick calls runAll = snapshotService.runAllSnapshots, which loops
runSnapshot per sport, which already carries the onGraded -> retention
hook. The scheduled path and the manual path are the SAME function. The
reason no cron cycle had been captured is simply that no slot has fired
since retention deployed (slots are 14/19/22/1/3 UTC; retention landed
~02:55). Induced proof follows the deploy.
PHASE 2 — archetype/team/opponent were permanently null because retention
persisted at GRADE time, before enrichment attaches them. Retention still
COLLECTS at grade time (the only moment the feature vector exists) but now
PERSISTS after enrichment, merging those three fields via
retentionService.mergeEnrichment. The merge is pure and fills ONLY those
three fields — features and every model output are grade-time values and
must never be rewritten by enrichment; a test asserts that. Unmatched rows
(refusals not in the enriched slate) keep nulls rather than guesses. The
empty-slate early return now persists too: a refusal-only slate is still
history worth keeping.
PHASE 3 — ZERO-WRITE ALARM. opsWatch.retentionZeroWriteAlarm pages at
missed-snapshot severity when a slot GRADED props but retention wrote
fewer rows than the slate (or nothing). runSnapshot now returns
retentionRows so the scheduler can evaluate it. Retention is best-effort
by design so it can never break a snapshot — which means a broken write is
silent by construction. This is the counterweight. A slot that graded
nothing never false-pages; an absent count reads as NOTHING and still
pages, distinct from a reported 0.
Suite 280/3349 green, build exit 0. Outcome stamping deliberately NOT
implemented (depends on the settlement fix).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Phase 3 complete. The insurance chain is proven end to end rather than
assumed: dump -> validated -> pushed off-box -> verified on the box ->
pulled back down -> rebuilt into a live database.
Pulled vyndr-20260720-051158.dump FROM the Storage Box (not the local
copy) with the in-session key through the pinned host key, never
bypassing StrictHostKeyChecking. Restored into scratch Postgres 17: 715
archive objects, 42 public tables, ledger_entries with all 27 columns and
real spot-checked rows.
ASSERTION PASSED: ledger_entries restored 645 == live 645 (target >= 645).
model_snapshots restored 100/100, so the retention store shipped yesterday
is covered by backups from day one.
Records the operational gotcha the restore surfaced: the dump is written
by pg_dump 17 (Supabase 17.6) and pg_restore 16 CANNOT read it —
'unsupported version (1.16) in file header'. The first attempt failed on
exactly this. Any DR runbook must use PG17+ tooling. Restoring into
vanilla Postgres also logs 12 ignored errors (Supabase roles/extensions
absent locally) which are harmless.
Scratch DB torn down, pulled copy deleted, both dumps still on the box,
nightly cron untouched. No private key material echoed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Records off-box as WORKING with the root cause (key was only in Hetzner's
project store, never in the box's authorized_keys — the box previously
offered an EMPTY auth list) and the proof: offbox_ok:true, and the file
independently VERIFIED on the box via rsync --list-only through the pinned
host key (vyndr-20260720-051158.dump, 833,917 bytes, 05:12:28 UTC,
byte-identical to the local dump).
Env truth captured from the run output: key is correctly base64-decoded,
destination has no leading-slash bug.
Hardening recorded: host key statically pinned (accept-new gone, missing
pin refuses the push), remote dir guaranteed, failed required push now
pages at urgent with offbox_ok:false while the exit code still tracks
on-box durability.
Flags the ONE outstanding acceptance item honestly: the round-trip restore
is NOT done, because the dev box cannot authenticate to the Storage Box
(the authorized key is Kev's, not the in-session keypair) and the
container has no Postgres server. Lists both unblocks and the assertion
target (>= 645).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Exit 0 from the backup script is deliberately tied to ON-BOX durability,
so it is not proof the off-box copy landed. GET /api/internal/backup/offbox
runs rsync --list-only against BACKUP_REMOTE using the SAME pinned
known_hosts as the push (checking never disabled) and returns the dumps
actually present, with size and timestamp — so off-box presence is a
verified fact rather than an inference from an exit code.
Needed because the dev box cannot authenticate to the Storage Box: the
authorized key installed there is Kev's ~/vyndr-backup-key, not the
keypair generated in-session, so independent verification has to run from
the container that does hold working credentials.
Suite 280/3338 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
PHASE 1 — HOST KEY STATICALLY PINNED. ssh-keyscan -p 23 returned an
ED25519 key whose fingerprint EQUALS the out-of-band value
SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM, so it is safe to pin.
scripts/storagebox_known_hosts now carries that verified line and ships to
the container (Dockerfile already COPYs scripts/). backup-db.sh uses
StrictHostKeyChecking=yes + UserKnownHostsFile=<pin> instead of
accept-new, which was trust-on-first-use and would have accepted an
impostor on the very first run. A missing pin file REFUSES the push rather
than silently falling back. Never weakened to accept-new/=no//dev/null —
a test asserts that on executable lines.
PHASE 1b — REMOTE DIR GUARANTEED. The box has only .ssh/, and rsyncing a
file into a missing parent either fails or silently writes the dump AS the
directory name — one file, overwritten nightly, reading as "backups exist"
while retaining exactly one. Uses rsync --mkpath when available, else an
explicit remote mkdir -p ahead of the push.
PHASE 2b — FAILED OFF-BOX PUSH IS NOW LOUD. Off-box is required, so the
failed-push path pages at "urgent" (was "low"/deferred) and the script
emits a machine-readable OFFBOX_OK=1/0/deferred that
POST /api/internal/backup/run surfaces as a distinct offbox_ok field.
Exit code deliberately still reflects ON-BOX durability — a good on-box
dump must not raise a false total-failure alarm. Surfacing the truth, not
manufacturing a failure.
No key material is echoed anywhere; only the PUBLIC host key is committed.
Suite 280/3338 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Top-of-file ground-truth block for orienting a fresh session.
Records what shipped tonight (probability layer revived 32/32, value
engine arc 1, grade-range work, backup durable on-box at 643/643,
model_snapshots retention live with 100 rows incl 36 refusals, ESPN
parser fix) WITH two honest qualifiers rather than a clean win: A still
does not emit in production so the A-RATED marketing hold stands, and EV
is overconfident (+62%/+61%/+56.9% captured, p_win clamps at 0.95) while
hero v2 already ranks on it.
OFF-BOX BACKUP recorded as NOT WORKING and deferred — never succeeded
once, every dump lives only on the Hetzner volume. Documents the two real
blockers fixed (missing base64 decode; container had rsync but no ssh
binary) and the remaining one: the Storage Box offers an EMPTY auth-method
list, which is an account refusing all auth rather than a wrong key.
Explicitly marks as UNVERIFIED that no Chrome/UI diagnostic was run — no
data on the SSH-support toggle, external reachability, project-vs-box key
scope, or any Hetzner outage — and lists those as untested hypotheses in
likelihood order rather than implying they were checked. The full
scratch-Postgres restore proof is recorded as still OWED.
Open items with status: settlement zero-pushes bug, ~28 props/day
unsettled, permanent model-version contamination (with the hard-cutoff
rule), A-grade unreachable, EV overconfidence, edge_pct broken scale, C4
CLV, consistency CV stopgap. Plus the three live credentials to rotate:
Storage Box password, VYNDR_INTERNAL_KEY (pasted in a transcript), and the
GitHub PAT still in .git/config.
Next queued: backtest harness (needs ~2wk history, currently holds one
night), settlement audit, opponent-strength sourcing (MLB solved via
statsapi pitching splits; NBA/WNBA open) behind the source-adapter
pattern, then the metrics engine gated on the harness.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Records model_snapshots as LIVE and verified capturing (100 rows over 2
cycles: MLB 14 graded/36 refused, WNBA 50 graded; features, grade_11,
p_win, ev_pct on 100% of graded rows).
First-ever refusal visibility: juiced_no_edge 18, rare_event_over_below_
line 13, insufficient_data 5 — the MLB gate refused 36 of 50 sides (72%),
now measurable for the first time.
Flags EV as OVERCONFIDENT and not fit to surface: first captured values
include +62.1%/+61%/+56.9%, which real markets do not offer. Cause is the
estimator clamping p_win at PROB_CEIL 0.95 off ~10 games. Hero v2 already
ranks on ev_pct, so it will pick the MOST overconfident read — calibration
must gate this before EV drives anything user-facing.
Logs the two settlement-correctness findings Kev asked to track (zero
pushes across 470 settled rows; ~28 props/day never settling) and the
ledger model-version contamination, with the rule that any backtest off
existing history must treat the 2026-07-19 fix boundary as a hard cutoff.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
The off-box push failed with 'rsync: Failed to exec ssh: No such file or
directory (2)'. The container had rsync and pg_dump from S62 but no ssh
binary, and rsync shells out to ssh for every remote transport. The dump
itself succeeded, so this failed AFTER a good backup and reads like a
network/auth problem when it is a missing package.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
RETENTION (Phase 2, priority zero). History starts compounding tonight.
migration 025 model_snapshots — APPLIED to prod. Append-only, one row per
graded prop PER SIDE PER CYCLE, with a unique index on
(snapshot_id, player_key, stat, line, side) so a retried cycle cannot
duplicate. RLS on, service-role writes only.
What it captures that the ledger never did:
- features jsonb — the model's INPUTS. Without these a backtest can only
grade our own homework; with them any future model can be replayed
against the exact conditions this one faced.
- REFUSALS (refused + refusal_reason). The ledger drops them, so a gate
refusing props that would have WON is invisible — unmeasurable lost
edge. Captured via a new onGraded hook in gradeSlateService that fires
with BOTH sides before any filtering.
- grade_11, the pre-collapse grade. The 4-letter map throws away the
entire live C-/C/C+/B- range.
- model_version + code_sha on every row. ledger_entries mixes pre/post-fix
grades with no marker and cannot be separated retroactively.
- p_win / ev_pct / fair_odds / takeable / value — none of which any
permanent store held.
Wiring: analyzeViaEngine1 attaches _features/_grade_11 (underscore =
internal); gradeSlateService fires onGraded then STRIPS them so they never
reach a cache or API payload; snapshotService builds rows and persists
best-effort. Retention reuses the LEDGER's dateET/gameIdFor helpers so
rows share the ledger's natural key exactly — otherwise the settle pass
could never join outcomes onto them. Rows are written BEFORE the empty-
slate early return: a slate that refused everything is exactly the case
worth recording.
CONTRACT HELD: retention is injectable and every path is caught. persist()
returns errors, never throws; a missing Supabase client is SKIPPED, not an
error. A retention failure can never break a snapshot.
BACKUP: backup-db.sh now accepts BACKUP_SSH_KEY as base64 (recommended —
survives env-var newline mangling, which is how injected SSH keys usually
break silently) OR raw PEM, detected by decoding and looking for the PEM
header. Verified both forms detect correctly against a real generated key.
Suite 279/3325 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
(a) WHAT REPLAYABLE HISTORY EXISTS — the headline is confirmed and worse
than "6 days".
ledger_entries is the ONLY store of model history in the database: 640
public rows, 6 distinct game days (Jul 11/12/16/17/18/19 — 13/14/15 are
missing entirely), 2 sports, 215 players, 470 settled, 465 settled WITH
odds. Every other candidate is 0 rows: grade_history, line_snapshots,
historical_props, closing_lines, resolution_results, accuracy_tracking,
model_predictions_extended, engine1_weights, prediction_registry and ~30
more. A data warehouse was designed and never filled. Redis holds no
history either (latest/previous at 24h TTL; the outcomes log carries no
odds/confidence/projection).
The blocking gap is not the day count, it is that NO MODEL INPUTS ARE
STORED ANYWHERE. No feature vectors, so we can score the grades we
emitted but cannot ask whether a different model would have done better —
which is the only question a harness exists to answer, and the exact gate
the metrics-engine north star requires. Also missing: p_win/ev_pct/
fair_odds (born tonight, on no column), grade_11 (only the 4-letter
collapse is stored, so the entire live C-/C/C+/B- range is unrecoverable),
and any model_version, so pre- and post-fix rows are already silently
mixed in one table. CLV remains unusable (C4). Settlement gaps surfaced
too: Jul 17 MLB 86 graded/57 settled, Jul 18 103/75, and 0 pushes across
470 settled rows — both feed the settlement-correctness audit.
Verdict: we cannot meaningfully backtest yet. Retention is priority zero;
every night without it is history we can never recover.
(b) DESIGN PROPOSAL — model_snapshots in Postgres (not Redis, which is
what lost us history twice). One append-only row per graded prop PER
CYCLE, capturing market values, model output, outcome (stamped later by
the settle pass), and critically a `features` JSONB — the counterfactual
enabler. Carries model_version + code_sha so eras never mix, grade_11 so
resolution is not thrown away, and refused/refusal_reason because
refusals are training data the ledger currently discards entirely.
Written from snapshotService (the existing chokepoint), best-effort so a
retention failure can never break a snapshot.
Volume: ~800 rows/day ~ 292k/year, ~300-600MB/yr of features, which would
exceed the Supabase free tier alone — so the proposal keeps full features
90 days and scalars forever. Three open questions for Kev before building:
the 90-day policy, whether to backfill the 640 existing rows as
scalars-only with explicit null features, and confirming we store
refusals.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
The real backup run failed: pg_dump could not write to /app/backups —
'Permission denied'. Cause: the container runs as the non-root 'vyndr'
user (Dockerfile USER vyndr) and the Coolify-mounted volume is
root-owned, so the mount is present but unwritable.
- Dockerfile now creates AND chowns /app/backups to vyndr alongside the
existing /app/data + /app/.pm2 line. Docker seeds ownership into a
NAMED volume on first creation, so this fixes it for a fresh volume;
a host bind-mount still needs a host-side chown, which is why the
next change exists.
- GET /api/internal/backup/verify now reports process uid/gid,
backup_dir_writable and the access errno, so the exact chown target is
observable instead of guessed. A mounted-but-unwritable volume reads as
'configured' everywhere else — this makes it loud.
Suite 278/3310 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
BACKUP_DIR is now a persistent volume (/app/backups), so the dump already
survives redeploys — the container-ephemeral risk that made this urgent is
closed. Storage Box SSH auth is not sorted yet, so the off-box push is
explicitly DEFERRED rather than failing:
- gated on BACKUP_OFFBOX=1 (plus BACKUP_REMOTE and BACKUP_SSH_KEY); until
then the script logs "off-box push DEFERRED" and exits clean.
- if an enabled push DOES fail, it is a LOW-priority "deferred" notice, not
a failure — the durable on-box dump succeeded, and calling that an
incident would train us to ignore backup alerts.
Adds the read-back check, because a backup nobody has read is a hope:
countRowsInDump() runs `pg_restore --data-only --table=X -f -` and counts
the rows between `FROM stdin;` and the terminating `\.`, proving the
archive CONTAINS the data rather than merely parsing. Needs no Postgres
server, so it runs inside the API container. Validated against a real
pg_dump from a scratch Postgres: counted exactly 604 rows.
GET /api/internal/backup/verify exposes it (newest dump in BACKUP_DIR,
size, table, rows_in_dump). Unit tests inject spawn/fs so CI needs neither
docker nor pg_restore.
Suite 278/3310 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
NORTH STAR (design philosophy, not built): VYNDR measures players by
MODERN FUNCTION, not legacy label — the principle already under the
archetype system, from Rashad Phillips' Basketball Position Metric. The
rule: every proprietary metric is baselined against the player's
functional ARCHETYPE's CURRENT-SEASON behavior, never the position's
inherited standard. The edge is that the market often prices today's
players against yesterday's baselines, so archetype-vs-position baseline
disagreement is a repeatable mispricing. Generalizes across sports. Moat
= proprietary metrics x current-game calibration x our private outcome
data. Metrics ship as VALIDATED FAMILIES: hypothesis, flagged build,
backtest, ship-or-delete with the negative result written down. Nothing
is real until the harness proves it predicts better.
SOURCING SCOPE (report, no code): MLB opponent strength IS derivable from
statsapi, verified live — one free call returns all 30 teams' pitching
splits (era/whip/avg/slg/ops/homeRuns/strikeOuts/HR9), which beats the
ESPN field we were reaching for because it is STAT-SPECIFIC, exactly what
opp_rank_stat wants. NBA/WNBA cannot use ESPN (its team endpoint carries
only a team's own stats, no defensive rating or pace); options are
stats.nba.com dashboards, deriving allowed-points from scoreboard finals
we already fetch, or API-Sports. API-Sports is a fallback tier at best —
100/day will not survive per-team-per-day. ESPN stays last, always behind
an adapter.
Proposed the SOURCE-ADAPTER pattern: one interface per feed, config-driven
primary+fallback per (sport x capability), normalized output so vendor
quirks stay in adapters, fallback announced rather than silent, sources
with zero callers deleted rather than left as corpses, and a health check
that PAGES when a source returns empty or broken — where EMPTY IS A
FAILURE. Tonight's crash (captured 0 / errored 15) and the months-null
opp_rank_stat are both exactly what that check exists to catch.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Closing the backup for real. Three changes, each fixing something that
would have made the Storage Box target fail or silently rot.
1. SSH KEY COMES FROM ENV, not from the container. Generating a keypair
inside the API container was the obvious move and it is wrong: the
container filesystem is ephemeral, so the key dies on the next
redeploy and the off-box push starts failing silently. backup-db.sh
now reads BACKUP_SSH_KEY (a Coolify secret), writes it to a 0600 temp
file per run, and removes it on exit via trap.
2. PORT 23, verified live. Hetzner Storage Box runs full OpenSSH on 23;
port 22 answers with mod_sftp (SFTP only). Banner-checked both against
u635423.your-storagebox.de. rsync now uses
-e "ssh -p ${BACKUP_SSH_PORT:-23} ... -i <key>"; the old invocation had
no -e at all and would have gone to 22.
3. OFF-BOX PUSH IS NIGHTLY, not Sundays-only. A weekly push meant up to
six days of dumps existed ONLY inside an ephemeral container, which is
the same as not existing. Alert copy updated to say exactly that when
the push fails or is skipped.
Also adds POST /api/internal/backup/run (internal-key gated) so a real
backup can be TRIGGERED and OBSERVED — it returns exit code, duration,
output tail, and whether the remote + ssh key are configured. The backup
can only run where SUPABASE_DB_URL and the Supabase route live (this
container), and there was no way to fire or inspect it without a shell.
Connectivity established this session: Storage Box reachable from the dev
box on 22/23; Supabase :5432 NOT reachable from WSL2 (so the dump must
run in-container, as designed); docker IS available locally, so the
restore-verify can run against a scratch Postgres using the real dump.
Suite 278/3305 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Ran the manual regrade with the internal key (thanks). Results are mixed
and the honest half matters more.
CONFIRMED WORKING — the probability layer is fully alive in production.
After POST /api/internal/snapshot/{mlb,wnba}: p_win, ev_pct, model_odds
and value are present on 32/32 live grades (mlb 7/7, wnba 25/25), up from
0/8 before. That fix is done.
NOT WORKING — the grade-range half did not land, and I am not going to
claim it did. The live distribution is unchanged (wnba B17/C8 before AND
after; mlb B4/C3), no A, no D, same four confidence values. Diagnosis:
matchup_grade is 0/25 on the live board, i.e. opp_rank_stat is still
null, so engine1's +/-1.0 opponent factor still never fires and the
ceiling is still +3.0 against the +4.5 an A requires.
Two distinct causes, both verified against the live ESPN feed:
1. refreshTeamStats CRASHED on every team — "buckets is not iterable",
captured 0 / errored 15. ESPN's current shape is results.stats =
an OBJECT with categories[], not an array. The old parser did for...of
on it. This was invisible until S63 gave the function its first
production caller. FIXED here (now captured 15 / errored 0) with a
regression test covering the current shape, the legacy array shape,
and empty payloads.
2. Even parsed correctly, the endpoint does not carry a
defensive-strength metric at all: defensive_rating, opponent_ppg,
pace and opponent_fg_pct all normalize to null — it returns only a
team's OWN stats. So defensive_rank_normalized cannot be computed and
opp_rank_stat remains underivable from this source. A test documents
the gap and will fail if that ever changes.
Consequence: A STILL DOES NOT EMIT, so the A-RATED marketing hold STAYS.
Reviving the opponent factor needs a different derivation (opponent
points allowed from scoreboard/schedule, or a different ESPN endpoint) —
logged as the concrete next item, not hand-waved as done.
Suite 278/3305 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
FOUNDATION-FIRST re-order, phase 1 (tooling + safety).
BACKUP (highest-severity open item) — INSTALLED, not re-proven.
src/backupScheduler.js runs scripts/backup-db.sh nightly from inside the
API container, armed at boot in server.js. The container already has
SUPABASE_DB_URL, pg_dump and the Supabase route, so deploy == installed:
no host crontab, no Coolify click. Arming is deliberately opt-OUT (armed
whenever SUPABASE_DB_URL exists; BACKUP_CRON=0 kills it) because the S62
design was opt-in and nobody ever opted in — the DB went unbacked every
night for weeks. A failed run pages high-priority ntfy; silence is the
danger with backups.
Durability is the one part still needing a human: the container FS is
ephemeral, so a dump dies on redeploy unless BACKUP_REMOTE (off-box
rsync) or BACKUP_DIR (persistent volume) is set. The scheduler detects
that and pages a WARNING at boot rather than letting an undurable backup
read as "backed up". Runbook rewritten to lead with the code path.
MANUAL REGRADE TRIGGER — scripts/run-snapshot.js, runnable via
docker exec with no VYNDR_INTERNAL_KEY and no new HTTP surface. Runs the
SAME snapshotService.runSnapshot the cron runs (including the team-stats
refresh that powers opp_rank_stat), supports `all` and `--settle`, and
prints the grade/confidence distribution plus p_win/ev_pct presence —
which is the thing you actually want when verifying a grading change.
ACCESS BLOCKER, logged honestly in specs/model-train.md: there is no
VYNDR_INTERNAL_KEY in the local .env and SSH to the box times out from
WSL2, so I can neither curl the internal endpoints (which already exist
from S45) nor docker exec. The trigger is built and correct but only Kev
can run it until a key or SSH access exists. This is the highest-leverage
unblock for phases 2 and 3, which both need on-demand regrade+settle to
verify anything.
Also logged the standing cautions: CLV ledger stays private until
backtest-proven; "self-improving model" is unsupported marketing until
the loop closes; the engine is MLB/WNBA-calibrated and NFL/NBA/soccer
need their own calibration before the hub grades them (scaling gate).
Suite 277/3300 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Not new work — logging so neither gets lost.
U-deg part 2: edge_pct is on a broken scale and it is the number users
actually see. Live: edge_pct 100 on a single-digit-edge prop; ledger-wide
311/604 rows (51.5%) exceed the frontend's sane cap of 40, 39 exceed 100,
worst 620. Mapped the consumers, and the split is the whole problem: 13
frontend files + deskShowcase/contentTemplate/parlayScan/tierGating read
the BROKEN edge_pct, and ledgerService:199 persists it to the
column of an append-only table right now. NOTHING on the frontend reads
ev_pct; only heroPropService does. Noted that S-b (rank board on EV) is
the real remedy and should be done as one piece with the scale fix, and
that EDGE_BOARD_SANE_MAX is damage control that nulls half the board.
Dispersion classifier: MIN_MEAN=4 is the honest stopgap; it leaves a
+/-1.0 dead for MLB low-count stats. The scale-free fix is variance/mean
vs the Poisson baseline of 1.0. Logged with its explicit validation bar —
backtest harness first (still does not exist), replay settled outcomes,
show no tier degradation, report before flipping, env-gate it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
POST /api/analyze/prop on prod returns p_win 0.523, ev_pct -10.4,
model_odds -109, confidence_basis grade_band, value false — every one of
which was absent on 100% of grades before this change. The value triplet
is whole (book -140 / fair -125 / model -109) and correctly refuses to
call a -140 price value when the model gives it 52.3%.
A-emission still pending the 01:00 UTC snapshot (opp_rank_stat populates
only when refreshTeamStats runs in a snapshot). MARKETING HOLD on A-RATED
copy stays until that passes. edge_pct scale remains broken (U-deg pt 2).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Folds re-sequenced steps 1+2 into one change (Kev's call): same bug
family — features wired to sources that return null.
THE PROBABILITY LAYER WAS DEAD IN PRODUCTION. p_win/ev_pct/kelly/
model_odds/value were absent on 0/8 live grades because
gameLogService.getGameLogs returns null for MLB by construction and
depends on the offline Python service for NBA/WNBA, so meta.gameLogs was
[] for every sport. This was the S46 bug in a second location — that fix
gave featureCache an MLB branch (why grades still worked) but never the
estimator. featureCache.getStatRows now supplies normalized rows
([{date,[statType]:v}], most-recent-first) for every sport, feeding the
estimator AND consistency AND game_count_in_7d from one fetch.
VERIFIED on real props: p_win 25/25 WNBA, 8/8 MLB (was 0).
GRADE RANGE, ON MERIT — never by rescaling (permanent founder ruling:
minting A's without new information is a relabelled B sold as an A and
corrupts an append-only ledger).
- refreshTeamStats wired into runSnapshot — it had ZERO production
callers, so opp_rank_stat was permanently null and a +/-1.0 factor
could never fire. Test-env no-op (opsNotify precedent).
- L20 made SYMMETRIC: both branches were delta +1.0, so the season
baseline could only ever ADD. No negative path was a structural reason
D was unreachable. New l20_contradicts_* carries -1.0.
- game_count_in_7d derived from real logged dates (heavy_workload_7d).
- NOT wired, deliberately, with reasons inline: teamId (no team_id
column; getFeatures reads it top-level; factor also needs a starter-id
list) and season_type (ESPN 2 = REGULAR season; threading it raw would
fire veteran_in_playoffs in July). Dead code dressed as a fix is the
thing we are removing, not adding.
CALIBRATION GUARD (found by verifying, not assuming): consistency CV is
NBA-tuned; for a Poisson-ish stat cv ~ 1/sqrt(mean), so any stat with
mean < 4 auto-classifies boom_bust. First verification run showed 8/8 MLB
props boom_bust — a blanket -1.0 that dropped the board to all-C. Floored
at CONSISTENCY_MIN_MEAN=4 -> 'unknown' below. Absent beats wrong. MLB
low-count stats therefore still get no consistency factor: honest, not
fixed. Scale-free index-of-dispersion classifier is the open follow-up.
CONFIDENCE IS NOT A PROBABILITY: payloads carry confidence_basis:
'grade_band'. Corrected mlb-grade-degradation.md — its "25/25
grade<->confidence agreement" is a TAUTOLOGY (confidence is derived FROM
the letter, so it would report 25/25 even if every grade were wrong), not
a validation. Removed dead mlbGrader.js (referenced only by its own test)
and the stale computeFeatures comment claiming a penalty that never ran.
VERIFICATION (scripts/verify-grade-range.js, real props/logs/engine):
WNBA 25 props B 68%->32%, C 32%->64%, D 0->1 (4%); 11-step spread went
from 2 steps to 5 (C/C+/B-/D). The D is earned: Angel Reese assists o2.5,
p_win 0.365. Nothing flooded — grades got HARDER. A did not emit locally
because opp_rank_stat needs the Redis cache only prod populates (local
ceiling +3.0 vs the +4.5 A needs); reachability is proven arithmetically
and locked in tests. Prod A-emission is the outstanding fingerprint.
MARKETING HOLD: "A-RATED" (AccuracyBadge, TopSignals) is unsupported
until that fingerprint. Confirmed honest fallbacks render today —
/api/ledger/accuracy has B and C buckets only, so the badge shows
"MODEL · 63% HIT" and TopSignals self-hides. Nothing fabricated ships.
Suite 276/3286 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Completes the diagnosis. Report only; no grade logic or thresholds changed.
The grade is an integer index (GRADE_SCALE, NEUTRAL_INDEX 3) moved by a
flat sum of +/-1.0 and +/-0.5 factor deltas, then clamped and rounded.
grade_thresholds.json is NOT an input mapper in the JS path — engine1
reads it BACKWARDS, taking the letter the index already produced and
looking up that band's midpoint to manufacture `confidence`. So
confidence is a cosmetic re-encoding of the letter: zero information
beyond it, and it can never disagree with it. There is no
data-sufficiency penalty in the live path (the one CLAUDE.md describes is
in mlbGrader.js, which is dead code).
Six of thirteen factors are wired to features nothing populates —
verified: refreshTeamStats has ZERO production callers (so opp_rank_stat
is permanently null, killing a +/-1.0), teamId/season_type/
game_count_in_7d are never passed (gameContext is built as {home_away}
and nothing else), and MLB consistency starves on the same dead
gameLogService path as Finding 2. Also verified: BOTH l20 branches are
delta +1.0 — there is no negative L20 contribution at all.
Arithmetic: an A needs sum >= +4.5; the live maximum is +3.0 (+2.0 on a
back-to-back, and MLB rest_days is 0 most days). D needs <= -1.51; the
live minimum is -1.5 and Math.round(1.5)=2, so it misses by one rounding
tick. Reachable band is index 2..6 = {C-,C,C+,B-,B}, which the adapter's
FOUR_LETTER_MAP (a 3->1 collapse) renders as exactly {C,B} — the observed
output, derived from first principles. Reachable confidences {42,47,52,
57,63} match the live values {47,52,57,63} exactly; C- is truncated by
gradeSlateService keeping the higher-confidence side.
mlb-grade-degradation.md's "25/25 grade<->confidence agreement" is a
TAUTOLOGY, not a validation — confidence is derived from the letter, so
it would report 25/25 even if every grade were wrong.
Recommends feeding the starving factors (restores A/D on merit) and
explicitly REJECTS re-scaling thresholds, which would mint A's without
adding information — every "A" would be a relabelled B.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Kev's call: investigate the B/C grade collapse before building. Report
only — no grade logic, thresholds, or engine code touched.
FINDING 1 — the collapse is real, live and structural. Across 604 ledger
rows and both sports the engine has emitted exactly TWO grades (B, C) and
NINE confidence values (63/57/55/52/47/45/35/25/20), ceiling 63. Still
true today on both sports. Confidence does NOT determine the letter:
conf 45 -> B while 47 and 52 -> C (non-monotonic), so the surfaced
confidence is not the quantity the letter came from. Edge scale still
broken: 311/604 rows exceed the frontend's sane cap of 40, 39 exceed 100,
worst 620.
FINDING 2 (bigger) — the entire probability layer is DEAD in production.
Live /api/snapshot/mlb: p_win, kelly, ev_pct, model_odds and value are
absent on 0/8 grades, while alt_lines (Desk-gated) IS present 8/8 —
proving nothing is tier-stripped, they are simply never computed.
Root cause: gameLogService.pythonPath returns null for MLB by
construction and the Python service is offline for NBA/WNBA, so
meta.gameLogs is [] for every sport; estimateProbability returns
p_over null; every field guarded by `if (pWin != null)` is skipped.
This is the S46 bug in a second location — that fix added an MLB branch
to featureCache.gameLogFeatures (which is why grades/projections still
work) but never to the estimator path.
Consequences: EV — the Model Train's whole ranking signal — has never
been computed on a live prop. Hero v2 matches nothing and always falls
through to the recent-read fallback (live /api/hero-prop returns
is_recent:true). Quarter-Kelly, sold on the pricing page and listed BUILT
in PROMISE-AUDIT.md, never runs. The value triplet is a duet live.
Recommend re-sequencing: revive the probability layer BEFORE G-a and
C-led (C-led would persist a column of nulls; G-a's EV_FLEX_THRESHOLD
would gate on a permanently-null value — Kev's EV_FLEX_ENFORCE=0 ruling
accidentally prevented an outage). featureCache:206-226 already has both
adapter branches and is the template.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
REPORT-FIRST per the arc order. G-a is HELD — the data changes the
recommended dials. No engine code touched.
Replayed against live ledger_entries (576 rows, 6 game days, 470 settled)
because the "30 days of stored snapshots" does not exist: snapshot Redis
keys are latest/previous only at 24h TTL, and no backtest harness exists
anywhere in the repo.
Findings that change the plan:
- The -400 floor shipped this morning was the whole win: past -400 hit
80.3% against an 86.9% breakeven = -13.29u / -7.7% ROI on 173 settled.
- Arc 2's incremental cut over the live gate is ~11 props in 6 days. The
only material change is gating the flex band behind 2x EV.
- The flex band (-161..-250) is our BEST band (+2.2% ROI, n=70) and the
takeable band is flat (-0.3%, n=209) — the opposite of the assumption
behind EDGE_FLEX_WALL. Recommend shipping the knob with enforcement
OFF until EV is persisted and measured.
- ev_pct/p_win are on NO ledger row, so the EV half of the gate cannot be
replayed at all. C-led (persist EV) is now the highest-leverage item.
- Confidence is monotonic but understates hit rate by ~20-25 points, and
the entire public ledger contains only B and C grades — zero A/A+.
That breaks hero v2 (isAB) and undermines "A-RATED" copy. Escalated.
- L-a answered: alt_lines carry NO odds and the feed has no alternate
markets. L-b is blocked on a data source, not engine work.
- C-led needs no odds backfill (locked_odds 99.1% populated).
- U-deg: the projection==0 leak is already closed (0 since 07-18).
- C4 confirmed in data (359/376 MLB closes == the lock). Stays suppressed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Arc 1 (7a925f4) shipped without a spec and without a STATE.md entry — the
plan lived only in a session context that was lost. Both written from the
code on disk, not from memory.
- STATE.md: header was stale at a8e383e; now 7a925f4 (pushed, NOT yet
deploy-fingerprinted). New top section records de-vig + EV + takeable/
value gates + hero v2 + the value triplet, the real config values, and
the 274/3289 -> 276/3306 test baseline.
- specs/model-train.md (NEW, per CLAUDE.md rule #1): every knob and its
ACTUAL default (TAKEABLE_ODDS_CEILING -160, TAKEABLE_ODDS_MAX +200,
VALUE_EV_THRESHOLD 2, JUICE_ODDS_FLOOR -400, RARE_EVENT_LINE_MAX 0.5);
the gate AS BUILT (flat price-only refusal at -400 pre-feature, NOT
edge-aware, no -250 wall; the -160..+200 band never refuses a grade and
is enforced on the hero alone); the triplet + hero v2; and a checklist
of what is NOT built.
Recorded honestly rather than assumed: EDGE_FLEX_WALL, HARD_JUICE_WALL,
LADDER_ODDS_MAX and MIN_RUNG_PROBABILITY do not exist in the codebase; no
frontend reads ev_pct/fair_odds/model_odds/book_odds/suppressed_reason yet;
the value fields are ungated to free tier; a second EV implementation
(processing/EVCalculator.js) still coexists with devig.evPct.
No engine code touched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
Steps 1-6 — make "real opportunities at takeable prices" the engine, not a filter.
1. DE-VIG (src/utils/devig.js): two-way multiplicative de-vig strips the vig and
returns fair prob + fair price per side + the overround. One side missing →
fair UNAVAILABLE (null), never faked. Method noted in code + the `devig_method`
field.
2. EV (devig.evPct): ev_pct = model prob × decimal − 1 at the graded side's
ACTUAL price. This is the ranking signal now, replacing raw |model−consensus|.
3. TAKEABLE gate (src/config/valueEngine.js, TAKEABLE_ODDS_CEILING −160 .. +200,
env-tunable): promoted surfaces only (hero/featured/alerts). The full board
still shows everything; Parlay Lab exempt; JUICE_ODDS_FLOOR (−400) stays the
absolute backstop underneath. Strict null-guard (Number(null)===0 would have
made a missing price "takeable").
4. VALUE flag: passes BOTH gates (takeable AND ev_pct ≥ VALUE_EV_THRESHOLD).
Grade = read quality; value = the price pays you. Shipped in payloads.
5. HERO v2 (heroPropService): highest ev_pct among takeable A/B reads — a huge
gap on a −900 line is trivia, not an opportunity.
6. VALUE TRIPLET: book_odds · fair_odds · model_odds on every read (snapshot,
hero, scan — they all spread the grade). Handoff documents the fields; the
rendering is Session-2 Design's job.
All wired in analyzeViaEngine1's existing p_win/kelly block (real quantile
probability × real book odds, or nothing). 33 new tests; suite 276/3306 green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Follow-up to the rare-event under fix — the whitelist (doubles/triples/HR/SB)
was fragile: the same juiced-under problem exists for steals, blocks, and any
other low-frequency market, and a new stat would slip through.
The real signal is the book's own price. The doubles unders were priced -625 to
-1100 — laying 6-11x to win 1x on an ~82% event, with no value the model could
recover. So the PRIMARY guard is now stat/sport-agnostic: analyzeViaEngine1
refuses any read whose graded-side odds are past the juice floor
(JUICE_ODDS_FLOOR, default -400, env-tunable). That catches every version of
this — steals, blocks, anything — and it also keeps the public record honest
(those -800 "wins" hit ~82% of the time and would inflate the hit rate, the same
class as the projection-0 degradation).
The structural rare-event rules stay as the BACKUP for props with no odds
(list also expanded cross-sport: + steals, blocks). Normal + longshot prices
(-110, -250, +600) are preserved. 16 tests cover both layers.
Reported: the doubles projection was REAL per-player (not a fallback); the fix
is the price guard, not a bigger list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Betting-logic audit: the CONSENSUS-vs-MODEL board flooded with fake reads like
"DOUBLES u0.5 · MODEL 0.2 · +edge" — the juiced under side of rare counting-stat
markets (doubles/triples/HR/SB), which is never a takeable edge and violates the
no-unders-default doctrine.
Report finding (item 3/4): the doubles projection is REAL per-player, not a flat
fallback — 'doubles' maps to a real game-log field (MLB_LOG_FIELD doubles→
doubles) and the live values varied (0.03/0.16/0.2/0.22). So no projection-gate
refusal for fakeness; the problem is purely structural (a rare event's real
projection always sits below a 0.5 line, so the under always "wins").
Fix (config-driven — src/config/rareEventMarkets.js, tunable stat list + line
threshold):
- Grade layer (analyzeViaEngine1): a rare-event UNDER at ≤0.5 is always REFUSED
(grade null + suppressed flag/reason). A rare-event OVER at ≤0.5 is refused
UNLESS the model genuinely projects the event above the line — because a
0.2-over-0.5 carries the SAME |edge| as the suppressed under and would just
take its rank on the board. The over grades normally once projection > line.
- Board layer (marketBreadth.collectBreadth): drops null-model rows so a
suppressed/ungraded prop can't rank a "MODEL —" placeholder onto the board.
10 suppression tests + config locks; also fixed a settingsPage book assertion
left over from the ESPN→theScore swap. Suite 274/3289 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
SUPABASE_DB_URL is set in Coolify on the API service, and this WSL2 box can't
reach db.<ref>.supabase.co — so the backup runs INSIDE the API container, which
has the env + Supabase network. Made that real:
- Dockerfile: install postgresql-client (pg_dump/pg_restore) + rsync + bash in
the runner image.
- backup-db.sh: added an integrity fingerprint on every run — pg_restore --list
must parse the archive AND find ledger_entries, else the run FAILS + pages
(stronger than the size check; catches a corrupt/structureless dump).
- BACKUP-RUNBOOK.md: rewritten for the container-exec reality — host cron does
`docker exec <api> sh /app/scripts/backup-db.sh` (inherits env + network +
pg_dump), or a Coolify Scheduled Task. Full restore-fingerprint steps included.
MECHANISM FINGERPRINT (run locally, docker + pg16): seeded a ledger_entries
table (137 rows) → ran backup-db.sh (dump + validate: 22 archive objects,
ledger_entries present) → pg_restore into a scratch DB → 137 rows restored,
exact match. The dump/validate/restore path is proven end-to-end; it's the same
pg_dump/pg_restore that run in the container against Supabase.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1+2. Checkout price was CODE-gated (founder price only with a valid founder
code) — so "Claim a Founder Desk" would have charged the $44.99 standard
price, not the advertised $34.99. Now it's SEAT-gated: resolveCheckoutPrice()
attaches the founder price while founder seats remain (< FOUNDER_SEATS_TOTAL,
read from the SAME countFounderSeats() truth as the ClaimMeter), and flips to
standard at seat 100. createCheckoutSession uses it; the founderCode param is
kept for back-compat but no longer drives price. The meter flips to "SOLD
OUT" at capacity. When the count can't be verified we honor the advertised
founder price (never overcharge).
- Also hardened countFounderSeats to manual pagination (the for-await form
broke on non-async-iterable list mocks).
3. Tests: resolveCheckoutPrice at seat 0 → founder, seat 100 → standard, the
99/100 boundary, null-count → advertised founder price.
4. Grace: invoice.payment_failed now sets a 14-DAY grace (spans Stripe's Smart
Retry window) instead of 48h — a transient decline no longer revokes access
mid-retry. Access is revoked only when Stripe actually cancels
(customer.subscription.deleted keeps its 48h grace). Test updated.
Stripe + founders suites green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
scripts/backup-db.sh: nightly full-DB pg_dump via the direct connection string,
14-day local rotation, weekly off-box rsync copy, ntfy alert on any failure +
an undersized-dump guard (an empty dump is a silent failure). docs/BACKUP-
RUNBOOK.md: the ONE env var Kev must set (SUPABASE_DB_URL — the direct
db.<ref>.supabase.co:5432 URI, not the pooler), the cron line, the off-box
target (Hetzner Storage Box via rsync, simplest for a Hetzner box), and the
restore FINGERPRINT procedure (pg_restore into a scratch DB + count
ledger_entries — proves it's a real, restorable backup).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
023_security_hardening.sql:
- Item 1 (CRITICAL, advisor lint 0010): founder_pricing_seats view → recreate
with security_invoker=on so it respects RLS instead of running as definer.
(The founder counter no longer depends on it — item 0 uses Stripe directly.)
- Item 3: waitlist write hole — drop the always-true policies, anon may INSERT
only, update/delete/read via service role.
- Item 5: pin an explicit search_path on the flagged functions (lint 0011).
024_anon_revoke_discoverability.sql:
- Item 4: revoke anon SELECT on the advisor-named tables (accuracy_tracking,
bets, cascade_alerts, closing_lines, coach_profiles, daily_scan) + a
commented broad sweep. The frontend reads data via Express (service role),
never as anon, so this is safe. REVOKE/KEEP rationale documented in the file.
These need Kev to apply (no DB access from here); fingerprint = re-run the
Security Advisor and confirm the lints clear.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The billing portal is fully configured in Stripe (cancellations, plan
switching, invoice history) and the Express endpoint (POST /api/stripe/portal)
existed, but nothing in the UI linked to it. Added the Next proxy
(app/api/stripe/portal) and a "Manage billing →" button in the profile billing
section (paid tiers) that mints a portal session and redirects. Kev still
activates the hosted portal in the Stripe dashboard; this is the app-side link.
Dunning verification (item 6): cancel-on-exhaustion is correctly wired — Smart
Retries exhausting cancels the subscription → customer.subscription.deleted →
webhook sets a 48h grace → middleware/gracePeriod.checkGracePeriod downgrades
tier to free in both users + user_profiles after the grace expires. See the
report for one nuance (the 48h grace on the FIRST payment_failed is shorter than
Stripe's 2-week retry window — self-correcting via subscription.updated, but
worth a product decision).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ESPN BET is defunct — PENN/ESPN terminated the deal; PENN rebranded it to
theScore Bet (Dec 1 2025) and ESPN is now exclusive with DraftKings. Removed
the ESPN BET entries from the BookChip map (web/src/lib/books.js) and added
theScore Bet (mono TS, slug thescore) as the successor. Added 'thescore' to the
backend oddsNormalizer ALLOWED_BOOKS so the feed's lines are accepted; synced
the bookWordmark test list. The ESPN references in src/config/sports.js are
ESPN's STATS API (data provider, unrelated to the sportsbook) — left untouched.
Flagged in specs/design-reference/HANDOFF.md that the design mockups' BookChip
row still shows ESPN BET and needs the same one-swap on the next refresh.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The counter showed 1/100 from user_profiles (founder_pricing=true AND
subscription_status='active'), but the live Stripe account has ZERO
subscriptions of any status — the "1" is a comped/manually-tiered profile, not a
paying founder. A tier/founder_pricing field on a profile can be set without
ever paying, so it is not proof of a paid seat.
Now the count is Stripe's OWN truth: stripeService.countFounderSeats() counts
ACTIVE subscriptions on a founder price. The route reads that (cached 5 min);
null or any failure → hidden, never a number. A comped profile no longer counts
→ the honest number is 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The blog showed "Posts coming soon" live: the app reads process.cwd()/content
= web/content at runtime (that's where the old orphan lived and rendered), but
the 5 articles were committed to REPO-ROOT content/articles — which the
deployed app never reads. Moved them to web/content/articles (verified
getAllPosts finds all 5 from cwd=web) and deleted the orphan file
web/content/blog/line-movement-guide.mdx (the route already 301s). Test paths
updated to web/content/articles.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The pricing Desk showcase hardcoded an alt-line ladder (1.5 A +7.1% / 2.5 A+
+11.4% / 4.5 C -3.8%), QUARTER-KELLY 2.4%, and PARLAY φ 0.34 — a mocked demo
selling something we weren't proving.
- deskShowcaseService reads the pre-graded snapshot for a real A/B prop's
alt-line ladder (prefers the one with the most grade variation — the most
compelling real example). Edge per rung shows only when it's a plausible
market value; the inflated (model-line)/line artifact on small lines is
guarded to "—" rather than shown as a fake +91%.
- PARLAY φ is now REAL: the model's same-team correlation (0.34, mirroring the
frontend parlayMath team constant) computed for TWO REAL same-team legs,
named. No real same-team pair on the board → the tile hides, never an
invented number.
- QUARTER-KELLY tile is REMOVED: the snapshot has no odds, so a real
quarter-Kelly % can't be computed here — a fabricated 2.4% is worse than
nothing. Kelly stays a real in-app Desk feature; the showcase just doesn't
fake it.
- DeskShowcase is now a client component fetching /api/desk-showcase; when the
board has no real ladder the whole visuals column hides (real-or-hidden, same
law as the hero). The pitch copy is unchanged.
5 service tests. Change-affected suites green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The landing hero was a static Jokic "Example" card with a name-length pick and a
hardcoded A- 73% +6.2% fallback. Now it's deterministic and live:
- heroPropService.pickHeroProp reads the pre-graded snapshot and selects the
prop with the LARGEST |projection - line| gap among A/B grades (conviction,
not noise) — the read where VYNDR disagrees most with the market, the card
that makes a stranger argue. No curation, no grading (reads cache → no API
credits). GET /api/hero-prop (backend) + repointed Next proxy.
- The card shows the disagreement EXPLICITLY: the book's line vs VYNDR's model,
side by side (model in green), with the real grade timestamp ("Graded 2:14
PM"). The EXAMPLE chip is gone.
- Empty slate → the MOST RECENT real graded read (flagged "LATEST READ", real
date). Nothing cached → { available:false } and the card HIDES. No
hand-written fallback — the Jokic card is deleted. Survives a dead night: a
live rule shows tonight's real MLB read, never a phantom July NBA card.
7 service tests lock the rule (max-gap, A/B gate, projection/line required,
empty→recent, hidden, cross-sport). colorContract updated to the new
disagreement display. Change-affected suites green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Kev's call: the 30D accuracy surfaces must read TRUTH, not a cache that can't be
filtered. My earlier degraded-row exclusion only touched getModelAggregate
(Postgres); the public buckets/badge still read outcomeService (Redis outcome
log), which counts degraded projection-0 outcomes and has no field to filter on.
- /api/accuracy (AccuracyBadge) + /api/ledger/accuracy (buckets/ModelRecord)
now source from the clean Postgres ledger aggregate via new
ledgerService.getAccuracyView + accuracyBucketsFromAgg (model_value > 0
excludes degraded rows). Same response shapes → no frontend change. Redis
outcome log is now read by nothing public; it can age out or be rebuilt.
- BEAT CLOSE is a MEASURED-WRONG ZERO: captureClosing re-records the locked line
as the "closing" line, so clv is flat on the whole sample and beat_close reads
0% (comparing a number to itself). Full write-up: specs/audit-data/
clv-capture-broken.md (the fix belongs to C4). Until then, beat_close_pct +
clv_distribution are SUPPRESSED at the source (getModelAggregate, gated by
clvCaptureReliable() / CLV_CAPTURE_RELIABLE=1). Every public surface already
renders BEAT CLOSE only when non-null, so they all hide it now — no wrong zero
anywhere. HIT RATE (real) is unaffected.
Suite 271/3261 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The test locked the old draft/unwired state; item 8 intentionally published
the 5 articles to /blog with real dates. Updated the assertion to the new
published shape (title + real date + status: published). This was a
tests-before-commit miss on the item-8 push (3b7a1f5) — fixed forward.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The live /blog showed "How to Read Line Movement Like a Sharp" dated 2026-03-22
— an orphaned, uncommitted file that predates the product. Meanwhile 5
genuinely-real articles sat in content/articles/, unwired.
- blog.ts now reads content/articles (not the untracked content/blog). Maps
`excerpt` → description, skips `status: draft`, reads explicit `slug`.
- Added honest dates (2026-07-17, the real publish day) + flipped the 5
articles to `status: published`. Real content: how-vyndr-grades-a-prop,
why-our-misses-are-public, what-clv-is, how-streaks-lie, the-vyndr-originals.
- Retired the orphan: /blog/line-movement-guide 301s to /blog (next.config).
- Added a minimal, dependency-free, XSS-safe markdown renderer so headers/bold
render as HTML instead of literal "##". Each article keeps title/date/
read-time + the existing OG/JSON-LD metadata (affiliate-review ready). Richer
media (images/charts) is the separate design train.
Web build exit 0 (SSGs all 5 slugs); backend suite unaffected.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ClaimMeter rendered a fabricated "47 / 100 CLAIMED" (a hardcoded default; the
comment even said "Cosmetic conversion driver… Static here"). Now:
- GET /api/founders/count counts ONLY real paying founders — user_profiles
where founder_pricing = true AND subscription_status = 'active' (the
Stripe-webhook-synced mirror, so we never hammer the Stripe API). Cached 5
min in Redis on top of that.
- If the source is unavailable (Supabase unconfigured, query error, column not
migrated, client throws) the endpoint returns { available: false } and the
ClaimMeter renders NOTHING — counter and progress bar both hidden. We never
fall back to a number.
- A low real count is shown honestly (0 → "0 / 100"); the truth is the feature.
Next proxy at app/api/founders/count. 6 route tests cover real count, low
count, error/unconfigured/throw → hidden, and cache-hit. Suite 271/3260 green,
web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The product argued with itself: FAB/nav said "Scan", Free tier "5 scans",
ticker "MLB slate scanned" — while the Ledger says "MY READS". Swept every
user-visible surface to READ:
- BottomTabBar FAB + Nav link: 'Scan' → 'Read'
- Pricing free tier: '5 scans to try the model' → '5 reads …'
- StatStrip: 'Awaiting next scan' → 'Awaiting next read'
- Ticker badge + snapshotService event: tag 'SCAN' → 'READ',
'slate scanned' → 'slate read' (readSportOf parses BOTH old and new so
cached ticker items dedupe cleanly through the rollover)
- upgradePitch: 'You've scanned N parlays' / 'unlimited scans' → read/reads
Internal untouched (not user-visible): /api/scan routes, scan_count column,
scanning state, DemoScan/ScanIcon, scanlines CSS, the transitional SCAN
color-map key.
tests/unit/verbLaw.test.js is the enforcement: it fails on user-visible
scan/scanned/scans copy across web/src + src/services (skips comments). Suite
270/3254 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Work-order #4 closed. First post-deploy snapshot (2026-07-17 14:00:53 UTC)
re-graded with the fix. Before → after:
- projection==0: 9/25 → 0/25 (the nine now refuse)
- grade<->confidence: mismatch → 25/25 agree
- edge_pct: {20,60,100,140} cluster → 7 continuous values, all-positive projections
The lone remaining edge=100 is a REAL projection (Abreu hits, line 0.5, proj 1.0
over = 100% by (model-line)/line), not the old proj=0 degeneracy. Verified.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
scripts/validate-grade-fix.js checks the live MLB snapshot for the three
degradation signatures (projection=0, edge=100 cluster, grade/conf disagreement)
— run after the next 14:00 UTC regrade to fingerprint the fix. The finding doc
now records root causes, fixes (commits 888d103/9fc4edf), the blast-radius SQL
(box can't reach Supabase directly), and the shared-path note for NBA/WNBA.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The degraded grades (projection=0 → model_value=0) are already settled in the
append-only ledger and must NOT be deleted (Data Semantics law). But their
hit/miss is noise, not model skill — they never had a real projection. So
getModelAggregate now filters `.gt('model_value', 0)` on both the settled and
pending queries: the rows stay in ledger_entries, but leave the public hit_pct /
CLV / per-tier record. `.gt` also drops NULL model_value. Post-fix no such row
can be written (projection<=0 refuses), so this only sheds the historical set.
This is the functional form of the "marking" the work order asked for — the
degraded locks are effectively marked as non-counting without mutating history.
Test builder mocks gained `.gt`; a lock asserts the filter is applied to both
queries. Suite 269/3253 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1 board item — three grading bugs the phone audit surfaced, all in the
live Node grade path (engine1 + analyzeViaEngine1), fixed at the source.
1. PROJECTION=0 NOW REFUSES. projectionFor returned l5_avg even when it was 0
(finite, so the `== null` gate passed it) — 9/25 live grades graded on a
zero projection, producing a degenerate edge and a hollow grade. Now a
non-positive reference is not a projection: projectionFor skips it and falls
through to the next POSITIVE reference (l5 -> l20 -> per_90 -> xg); when none
is positive it returns null and the read REFUSES (insufficient_data). The
gate also gained an explicit `> 0` guard so the invariant is structural — a
grade can never be emitted with a non-positive projection. Fewer graded
props, honest.
2. EDGE_PCT. The formula was already (model - line) / line signed by direction
— Kev's intended semantics. The broken {20,60,100,140} cluster was the
proj=0 degeneracy ((line - 0)/line = 100%); with #1 those refuse, so the
fabricated 100s vanish and real edges flow. The main-line edge now reuses
the VALIDATED projection (edgePctFor accepts an optional ref) so edge and
the persisted projection can never diverge. Frontend |edge|>40 guard stays
as a safety net.
3. LETTER == THRESHOLD_TABLE(CONFIDENCE). engine1's hand-rolled
GRADE_TO_CONFIDENCE drifted a full sub-tier low (B -> 0.55, which the
canonical grade_thresholds.json calls B-) — the "B at 45%" the audit caught.
Now confidence is DERIVED from each grade's band MIDPOINT in
grade_thresholds.json (one source of truth, shared with the Python engine),
so applying the threshold table to any grade's displayed confidence resolves
back to the same letter. Proven for all 11 grades.
Regression locks: tests/unit/mlbGradeDegradation.test.js (14 tests). Backend
suite 269/3253 green, web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: read cards were huge, only 1-2 fit per screen. Compressed the
vertical spacing — article padding 16->12, header margin 8->5, name 15->14px,
ladder-rungs margin 10->8, book/date line 12->8.
Kept the archetype showDesc: it renders INLINE (same row as the badge), so it
adds zero vertical height — dropping it wouldn't help density and would break
the ds5 design lock ("the badge shows its one-line meaning where it leads").
Locked the density in vyndrParityQA (P2-10).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
DISPLAY FIX (shipped): the league leaderboard rendered raw snake_case
("stolen_bases U0.5", "earned_runs U2.5"). New canonical short-label lib
web/src/lib/statAbbrev.js (one source, CommonJS + unit-tested) maps stat_type
to SB/ER/TB/HR/K/PTS/… and ExploreHub routes through it. Unknown ids upper-case
their words so raw snake_case can never leak again.
FLAG (reported, NOT silently changed — per the audit's instruction): the "B at
45% confidence" is a BACKEND grading issue, diagnosed against live snapshot:
- 25/25 grades mismatch their own confidence vs grade_thresholds.json (B shown
at conf 55 = the B- band; a systematic one-sub-tier gap on every prop). The
surfaced `confidence` is not the probability that derived the letter (likely
the data-sufficiency penalty applied to display-only).
- 9/25 have projection=0 — the MLB feature path feeds 0 instead of refusing
(S58 insufficient_data), which also produces the P1-7 broken edge_pct.
Full write-up + do-not list: specs/audit-data/mlb-grade-degradation.md. NOT
re-lettering or shifting thresholds on the frontend — that would hide the bug.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: the Compare verdict read "the edge tonight tilts his way" for
Jokić vs Wembanyama — an NBA claim in July, when NBA has 0 games. The page is
sample/form data with no game resolution, so "tonight" can never be verified.
Reframed to "on current form" (the rows ARE L10 form) — always honest, in or
out of season.
Also fixed the cited dimensions: the verdict claimed "usage", but in the sample
Jokić's Usage% (29.1) is LOWER than Wemby's (31.0) — he wins scoring, boards,
and playmaking, not usage. Copy now matches the data.
Locks both P1-7 and P1-8 in vyndrParityQA.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit called the board 'mostly-empty'. Two causes, both now addressed:
1. Dead images (P0-2, already fixed) → the matchup/team chips render logos now.
2. Degraded edge data. Live snapshot edge_pct is on a broken scale (distinct
values 20/60/100/140 — not a market %), with projection=0 and confidence
35-55%. A real prop-market edge is single-digit, never past ~40%. Leading
the board with '+140%' fabricates a signal (Data Semantics Rule).
Fix: an edge whose |value| > 40 is treated as ABSENT at BOTH layers — the
data layer (flattenToEdgeBoard nulls it, so it can't RANK a fake +140% above a
real +8.4%) and the display (EdgeCell shows '—'). Board falls through to the
grade-rank tiebreak when edges are unreliable. Real edges (≤40) are untouched.
The root cause — edge_pct/projection/confidence degradation — is a BACKEND
grading issue (same family as the P2-9 '45% B' flag), reported separately; this
is the honest frontend guard, not a fix for the data.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: the same screen showed green 'SIGNAL LIVE · UPDATED 0s ago' AND
amber 'SYNC 46:03' — two components reading different fields. The Slate's
'UPDATED Xs ago' measured the CLIENT poll time (always ~0s, since it refetches
every 60s), while the app-bar clock honestly measured the pipeline's data age
(refreshed_at). '0s ago' claimed a freshness the data didn't have.
Now the Slate captures snap.refreshed_at from the snapshot it already fetches
(newest across sports) and 'UPDATED' shows THAT — the same source of truth as
the clock. The clock still owns the amber/STALE reaction; the strip just states
the honest data age. No more contradiction.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: CONSENSUS VS MODEL bled off the right edge ('2 BO…', 'MO…') and
STARTING-pitcher lines truncated ('2.64 E…'). The M4 lock only hid the DOCUMENT
scroll (html/body overflow-x) — content still clipped inside cards. Now contained:
- .breadth-row stacks (flex-direction:column) at <430px, each field on its own
line with overflow-wrap:anywhere — no bleed.
- the game-card starting-pitcher inner spans wrap + shrink (flexWrap + minWidth:0)
so name/ERA/archetype flow onto a second line instead of clipping.
- STRENGTHENED the lock: vyndrParityQA now asserts the CONTAINMENT patterns
(breadth-row stacks, pitcher spans wrap), not just document overflow.
MLB stat pills: the game-lines grid already scrolls-within-card (<640 M1); if
the audit still shows pill clipping elsewhere, it's a follow-up targeted pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: at 390px we still rendered the full desktop 3-row header (nav +
TOP MOVES ticker + SYNC line) eating ~20% of the viewport, and its height
clipped page titles under it (MY READS tabs, HEAD TO HEAD). Implemented Design's
mobile app bar <768px:
- New MobileSyncClock (extracted from HeartbeatBar) lives in the Nav's right
cluster — wall clock rests, amber/STALE reacts off the shared freshness tier.
- <768px: the ticker row (.nav-ticker) AND the whole heartbeat bar are hidden;
only the nav row shows (logo + clock + search). main padding-top → 62px and
the Slate sticky tabs → top:60px, so nothing clips under the bar.
- Locked in vyndrParityQA (P0-4): ticker+heartbeat hidden, nav clock shown,
paddings collapsed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Completes P0-3 across all three surfaces:
- CONSENSUS VS MODEL (MarketBreadth): dedupe by player+market so Ben Williamson's
alt-line variants show as ONE consensus row (no per-market cap — a consensus
table just shouldn't repeat a player).
- LEDGER cards: group by player+market via groupIntoLadders — Alec Bohm's
strikeout ladder (U1.6/U1.3/O1.5) is now ONE card with the rungs nested (each
its own side/line + tier-colored grade), not three separate cards.
- playerGrouping reads player OR player_name (ledger rows use player_name) —
regression-tested so the ledger doesn't silently empty.
The Alt Line Ladder shape is what /pricing already demos; the record now uses it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: leaderboard flooded with 9 consecutive identical 'stolen_bases U0.5
45% B' rows. New shared lib/playerGrouping (dedupeLeaders + groupIntoLadders,
name-key aware, 7 unit tests): ONE row per (player, market family) keeping the
best-ranked, then a per-market cap (4) so no single prop type floods the board.
Applied to ExploreHub. Ledger cards + Consensus grouping follow in P0-3b/c using
the same lib.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit: 78 ESPN <img> at naturalWidth:0, ZERO espncdn/mlbstatic requests
ever fire — the browser never attempts the fetch. CSP was NOT the cause (img-src
already allows a.espncdn.com + img.mlbstatic.com, confirmed in the live header).
Root cause = native loading="lazy" on the tiny entity <img>s never triggering on
device (the audit's suspect b). TeamLogo + PlayerAvatar now load="eager"
(+ decoding="async") — they're 10-40px core-visual entities, so eager is correct
and cheap, and it forces the fetch on parse regardless of size/viewport.
Browser-request confirmation is the audit's Chrome domain (I can't observe
network from server curls); eager semantics guarantee the fetch once the src
(which 200s server-side) is in the DOM.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Phone audit found LEDGER READ CARDS rendering B badges with BLUE borders + C
with AMBER right now in prod — GradePill (components/GradeCard.tsx) hardcoded the
OLD palette (rgba(74,158,255) blue-B, rgba(255,179,71) amber-C) for bg/border
while the text used the migrated token. Migrated bg/border to color-mix on the
grade token, so B renders neutral-white and C grey (matching the board).
- globals.css .grade-*-bg → token-derived color-mix (was raw blue/amber rgba).
- DELETED glow from .grade-glow-b/c/d (glow is A-tier ONLY, by law) — B/C/D keep
their token color, no text-shadow.
- Purged the last dead grade-blue #4A9EFF fallbacks (SoccerGradeResult, the
intelligence INFO dot).
- REGRESSION LOCK: vyndrParityQA fails if #4a9eff / rgba(74,158,255) reappears
anywhere in web/src, if GradePill hardcodes blue/amber rgba, or if grade-glow
B/C/D grow a text-shadow again.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Design colors every LIVE indicator green (active = green; the live-dots and
LIVE·Q3 labels are all green). The slate header's 'N LIVE' count was red
(--live #ff3b3b) — contradicts the one-meaning-per-color contract (red = miss/
error, which live isn't). Aligned to --g-a.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
PitcherArsenal's meta line showed team + 'vs OPP' as plain text; Rev 3 anchors
the pitcher's team and opponent with TeamChips (real logo / monogram). Combat
FightCard intentionally keeps fighter monograms (no photos — likeness rule).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Structural mobile rules become failing tests (M4 'test-lock what's lockable'):
the flat EDGE BOARD shows <768px and game cards are desktop-only; the board
renders TeamChips + tier GradeBadge + sign-colored hero edge% + the ranked
opacity ramp; document never scrolls sideways at 390px. Source assertions —
they lock the RULES, not the pixels (that's the master audit).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The one genuinely-new mobile screen. Design's mobile BOARD is a FLAT edge-ranked
list (all graded props across every game on one list, sorted by edge) — not the
desktop's game-grouped cards. Implemented to the drawing with REAL snapshot data:
- slateAdapter.flattenToEdgeBoard(cards) — pure transform of the assembled
GameCardData[] (grade→game join already done) into ranked rows, edge desc.
STRICT null edge sorts LAST (never 0-coerced to the top — Data Semantics Rule).
Threaded edge_pct through buildPlayerStripsFromProps (was dropped). 6 unit tests.
- MobileEdgeBoard component — Design's exact screen-01 rows: rank (green #1),
player + prop, matchup sub-line with TeamChips + live-dot, tier grade chip,
and the edge% as the one bold mono hero (green +, red −). Ranked opacity ramp
(1 → .55) + green inset border on the top reads. Breadth strip EDGES/AVG CLV/
GAMES — CLV honest '—' (per-slate CLV isn't computed; never fabricated).
- Slate: <768px renders the flat board, ≥768px keeps game cards (same data,
toggled by width). Ungraded slate still shows game cards on phones (no blank).
Built to Design's screen-01 drawing, VISUALLY UNVERIFIED at 390px — the core
mobile screen, top of the master-audit list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
GradeResultCard already carried Design's §7 grade-reveal structure (hero +
identity + 3-col grid + intel box) — the screen was built to the same contract
Design's mobile follows. Applied the Rev 3 gaps: team context now a TeamChip
(real logo/monogram), and the mobile grade hero → 74px (Design's exact mobile
size; M1a had 80). Desktop hero unchanged (116px).
Built to Design's mobile spec, VISUALLY UNVERIFIED at 390px.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Design Rev 3 anchors every matchup/context abbr with a 10-12px team tile. New
reusable TeamChip renders the real TeamLogo (licensed ESPN logo where it
resolves, team-colored monogram otherwise — the resolver is already built) at
that size + the abbr, sitting inside the row so it inherits the ranked opacity
ramp. First placement: StatStrip's player/team context (name → team-chip →
archetype). ROW-GRAMMAR identity-run test updated to the chip marker.
Remaining Rev 3 placements to thread TeamChip into (reusable, mechanical):
parlay legs, grade-shift header, pitcher "vs", /u recent-settled, other
matchup context lines. Game-card headers already carry TeamLogo (TeamLink).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Design Rev 3 drew the 9 legacy classifier marks (BRUSH #E0B84A, WHIFF #E86A6A,
CONNECTOR #9AB0C4, DISTRIBUTOR #7AB8D8, FASTBREAK #4AA0E8, FLEX #A08AC8,
HYBRID #C88AB0, SWITCH #C0B08A, SWITCHBOARD #90A0E8). Wired each to its own
Design mark + color (front lib/archetypes.js + backend archetypeService.js,
color-synced). These were the 9 NO-MARK backend keys — now none are on a
generic placeholder. They're classifier-side fallback renders (never
user-facing archetype names, per MANIFEST).
Re-imported Rev 3 package over specs/design-reference/ (83 glyph SVGs + MANIFEST
regenerated from the authoritative glyphDefs(); HANDOFF Rev 3 note).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Token alignment (M3.1) orphaned a few bare literals: PlayerAvatar's team-less
monogram accent still fell back to the retired B-blue #4A9EFF → Design neutral
#B8BCC8; the player/u OG + portrait billboards painted old text-0 #E8E8F0 →
Design #F0F0F0 (satori canvas needs literals).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The only surface still showing a raw book key: BookComparison rendered
{b.book} capitalized ('Betmgm'). Now uses the entity-layer BookChip (branded
wordmark + name), matching ledger/GameCard. No lowercase/raw book text anywhere.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replaced the generic reused shapes (triangle/star/plus/bolt…) with Design's
real per-archetype marks from the authoritative glyphDefs() (HANDOFF), and
aligned each archetype's color to Design's deduped palette — frontend
lib/archetypes.js + backend archetypeService.js kept in color-sync (the
cross-file test iterates the backend set). Badge test color expectations
updated to Design (BOMBER #FF9F45, CONDUCTOR #6C8CFF, ALPHA #7C5CFF,
FORTRESS #4C6FA5, …).
Scope + honesty:
- 29 non-combat archetypes wired to real marks + Design colors.
- COMBAT namespace (STRIKER/GRAPPLER/PRESSURE/COUNTER/FINISHER/GRINDER) left
untouched — it uses unicode CHAR glyphs + its own pinned colors + test
(FINISHER deliberately doesn't collide with the soccer FINISHER). Combat
could adopt Design's SVG marks in a follow-up.
- 39 Design marks are INERT (no classify() producer yet) — the 74 SVGs live in
specs/design-reference/assets/glyphs/; they light up when classify() expands.
- 9 backend archetypes have NO Design mark (BRUSH/CONNECTOR/DISTRIBUTOR/
FASTBREAK/FLEX/HYBRID/SWITCH/SWITCHBOARD/WHIFF) — kept on their generic glyph,
flagged for Design.
VISUALLY UNVERIFIED at 390px/desktop — archetype marks + colors on the audit list.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The share/OG card (a billboard) still painted grade B blue (#4A9EFF) on canvas;
Design's B is neutral-bright white. Canvas needs a literal, so #F0F0F0 direct.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reconcile finding: the live tokens had DRIFTED from Design's package (HANDOFF
"Tokens (exact)"). Per the standing conflict rule (Design specified exact hexes
→ rejected the drift → Design wins), aligned the whole palette:
- Surfaces: --bg-1 #0E0E16→#0E0E14, --bg-2 #15151F→#14141E; added --bg-deep
#0A0A10 + --hairline #101018 (Design's ramp).
- Text: --text-0 #E8E8F0→#F0F0F0, --text-1 #7A7A8E→#B8BCC8 (Design's secondary
is far brighter), --text-2 #4A4A5E→#707080, added --text-3 #4a4a58 micro.
- Borders: #1E1E2E→#1E1E2A, #2A2A3E→#2A2A38.
- GRADES (the big one): B blue #4A9EFF → neutral-bright WHITE #F0F0F0; C amber
#FFB347 → muted GREY #B8BCC8; D #FF5252 → #FF4757. The old blue/amber actually
violated DESIGN-SPEC v2's OWN "B neutral-bright, C muted" — this fixes a
long-standing drift, confirmed by Design's package. Amber stays its own token
(--amber); --warning decoupled to --amber; --miss → #FF4757.
- vyndrTokens.js GRADE_HEX mirror + the design-system test updated to match.
Big VISUAL change (grade color language), VISUALLY UNVERIFIED at 390px/desktop
— on the Chrome-audit list. Every surface inherits it, so it lands first.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Supersedes the prior partial import. From claude.ai/design project 370ba6df:
- HANDOFF.md (the entry point — exact tokens, glyph library, liftable
behaviors, laws)
- vyndr-mobile.html (19 mobile screens — closes all 10 previously-uncovered
surfaces: landing/pricing/streaks/ledger×2/empty/scan/player/team/explore +
combat/pitcher/archetypes/correlation), vyndr-system.html (desktop terminal),
support.js
- assets/glyphs/ — all 74 archetype marks as real SVGs + MANIFEST, GENERATED
from the authoritative glyphDefs() in the desktop file (currentColor, 24-grid,
duotone) rather than 74 fetches. These replace the generic star/plus
placeholders (the archetype family was never rendering as a distinct 74-mark
system).
- removed my interim MOBILE-SPEC.md distillation (the real 19-screen file +
HANDOFF supersede it).
Landing (desktop marketing) captured in-context — it's M2/M3-desktop scope;
raw file stays in the design project until that wave.
RECONCILE + build (token alignment, glyph lift, M1b surfaces, M3) follow.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Design's mobile app bar shows a wall clock, no SYNC/STALE readout. Building it
literally would drop the Ship-A staleness signal on the device most users are
on. Kev's ruling: Hybrid — the wall clock is the RESTING state (the stillness),
STALE/amber is the REACTION (the punctuation), driven by the SAME real signal
as desktop (refreshed_at vs expected_interval_s, thresholds 1.5×/3×). Never
silently stale on mobile. Design drew the happy path; we keep the failure state.
- web/src/lib/freshness.js — extracted the freshness tier as ONE shared source
(CommonJS, unit-tested); desktop HeartbeatBar + mobile clock both key off it,
so mobile can't silently disagree with desktop about staleness.
- LiveLayer: <768px hides SIGNAL LIVE + EKG + graded + the SYNC label and shows
a single right-aligned clock — ticking wall clock when calm, amber SYNC / red
STALE when the tier reacts. Resting dot is static (the clock is the pulse).
- Test-lock: freshness.test.js (tier thresholds, absent≠false-stale, no negative
age) + vyndrParityQA (mobile clock wired, one shared freshness source).
Built to Design's mobile spec, VISUALLY UNVERIFIED at 390px. NEXT shell
increments (still M1b): merge the clock into the logo row (Design's single-row
app bar), the breadth strip (EDGES/AVG CLV/GAMES — replaces the graded count
mobile lost here), and relocate the wire to the bottom of the board.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Design now provides mobile designs (Vyndr Mobile.dc.html, 5 screens at 390×844).
Captured as specs/design-reference/MOBILE-SPEC.md — the buildable distillation
(the raw dc-runtime HTML needs React to render; the design project holds the
verbatim source). I'm now IMPLEMENTING Design's mobile, not inferring it.
RECONCILE of the already-shipped M1a against Design's mobile:
- EKG hidden <768px — Design AGREES (mobile app bar carries no EKG). KEPT.
- .vbtn 44px tap target — Design AGREES (buttons are 44/46px). KEPT.
- .m-hero generic clamp — CONFLICT. Design uses CONTEXT-SPECIFIC hero sizes
(74px grade tier / 40px live tier / 24px /u stats), not one clamp, and
nothing consumed the class. REMOVED; the specific sizes land per-surface in M1b.
- Mobile HEADER shape — Design OVERRIDES: mobile is an app bar (logo + sync
clock) + a breadth strip + a bottom wire, NOT the collapsed heartbeat bar.
The EKG-drop is a compatible interim step; the full app-bar recomposition is
M1b (the board/shell surface). Flagged in the CSS + MOBILE-SPEC.
Built to Design's mobile spec, VISUALLY UNVERIFIED at 390px (Chrome audit is eyes).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
SETUP (required first step): overwrote specs/design-reference/ with the CURRENT
authoritative mockup ("Vyndr System.dc.html" from claude.ai/design, the version
with the combat card / pitcher identity / live grade-shift / correlation builder
/ FREE|PRO pricing / TRANSMISSION QUIET). 12 surfaces. Confirmed it has NO media
queries — desktop-only, so the 390px expression is a deliberate design decision,
not a shrink.
M1a — mobile foundation (built to spec, VISUALLY UNVERIFIED at 390px; WSL2↔Chrome
unreachable, the Chrome audit is the eyes):
- Header-zone collapse (the concrete audit finding: ticker + SIGNAL LIVE + STALE
stacking in ~110px). The decorative EKG is dropped <768px so the heartbeat
reads as ONE clean status line; signal pulse + ticker are the single animated
element (DESIGN-SPEC §4). LiveLayer gains .heartbeat-bar / .hb-ekg hooks.
- Primary CTA (.vbtn) meets the 44px touch target on mobile; dense `small`
buttons opt out (density is a feature).
- .m-hero mono-hero clamp for the one-figure-per-card grammar at 390px.
- M4 test-lock: vyndrParityQA asserts the header collapse, tap target, overflow
containment, and the honest "UNVERIFIED at 390px" label are all in source.
Per-surface stacking (M1b), billboards (M2), desktop parity (M3) continue.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Diagnosis (why 500/500 went unpaged): the only regular odds-api burner was
futuresService, which called axios DIRECTLY — bypassing the gateway, so it
never hit recordCall (the ONE place the WARN/BLOCK pager fires) and never
respected the 95% block. It only syncFromHeaders, which updated the counter's
number SILENTLY. oddsService (which does go through the gateway) only touches
odds-api when PropLine fails, so recordCall for odds-api effectively never ran.
Result: the counter could reach 100% with neither pager firing.
Fixes (a silent drain is now impossible, not just guarded):
- futuresService routes through gateway.fetch('odds-api', …) → counted, blocked
at 95%, and reserve-gated. Closes the raw-axios bypass.
- Reserve floor in the gateway: a DISCRETIONARY call (futures/soccer) passes
reserve=ODDS_API_RESERVE (default 50) and is refused while remaining <= reserve.
The ESSENTIAL MLB prop-backup passes no reserve and may spend to the 95% block.
→ a futures/soccer drain can NEVER starve MLB's backup path.
- quotaTracker.syncFromHeaders (the AUTHORITATIVE number) now fires the same
once-per-period WARN/BLOCK alert on a crossing — extracted fireThresholdAlert
shared with recordCall. The header-only drain now pages.
- POST /api/internal/quota/test-alert (internal-key) test-fires the pager
end-to-end so ntfy delivery is verifiable on demand.
Also (reality-corrected cadence): WNBA restored to the full grid. 2026-07-15
had two AFTERNOON WNBA games finished before the 22 UTC slot — 14 UTC (10am ET)
is the only slot early enough for a 1pm ET game's props, and on PropLine the
extra slots cost a rounding error. Soccer stays the only trimmed sport (the
real odds-api discipline). Assumption corrected by observed data.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Every ACTIVE sport was graded at all five MLB slots (14/19/22/1/3 UTC). Sports
post lines on different clocks, so that inheritance was wasteful both ways:
WNBA props aren't posted at 14:00 UTC (10am ET) → that slot always graded 0
(the audit's "wnba:0"); soccer odds come from the 500/MONTH odds-api key, so
five slots/day is a third of the budget for 1-2 matches.
New src/config/sportCadence.js is the single source of truth (config-over-
constants). Mapped from reality + quota headroom (PropLine 9k/day abundant,
odds-api 500/mo scarce):
mlb 14/19/22/1/3 intraday (full grid — games+props all day)
nba 14/19/22/1/3 intraday (in-season fits; off-season self-skips empty)
wnba 19/22/1 intraday (afternoon→evening ET; drops the 14/3 waste)
soccer 14/19 NO intraday (WC live; 2 lean odds-api reads, key-protected)
The scheduler still fires at HOURS_UTC and the missed-cron watchdog still
references MLB (which runs every grid hour) — each slot now grades only
sportsForHour(h), and only intradaySports() get the 20-min refresh. Every
sport's hours are kept a subset of the firing grid (a boot-time guard + a test
warn if that's ever violated). Retune a sport by editing one table row.
Adaptive, not constant: near-zero when a sport is quiet, protecting the scarce
odds-api quota from being drained by noon.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The "SIGNAL LIVE vs STALE 8h" contradiction was a field mismatch, not a dead
pipeline. updated_at is the grade-LOCK time (advances only on a full snapshot,
5×/day — grades never change in-game, so it is intentionally stable). The
SYNC badge measured the 20-min intraday cadence (expected_interval_s=1200)
against that 5×/day field → structurally guaranteed STALE between slots even
when intraday refreshes lines perfectly.
- snapshotService: full snapshot now seeds refreshed_at at lock time
- intradayRefresh already bumps refreshed_at every ~20 min (unchanged)
- /api/snapshot/summary + GET /:sport now expose refreshed_at (was written to
Redis but never serialized → no public liveness signal existed)
- LiveLayer SYNC badge measures freshness from refreshed_at (fallback updated_at)
Exposing refreshed_at also gives a public heartbeat probe: it advances every
intraday slot, so pipeline liveness is verifiable without container logs.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Wave 0 shipped NBA/WNBA grading off ESPN per-athlete gamelogs, but name→id
resolution went only through the v2 /search endpoint, which is unreliable at
the edges. Live probing surfaced the real coverage gap: search actually
resolves the right id for most names, but dual-league athletes (WNBA + NCAA —
e.g. Napheesa Collier, Brionna Jones) get a filters-only gamelog until a
`?season=` is supplied, so they silently returned insufficient_data despite a
full season of games.
Two fixes:
1. buildAthleteRosterIndex(sport) — aggregates every team roster for nba/wnba
into a complete { nameKey → {id, displayName, teamId} } map (canonical
accent-folded keys via playerName.nameKey). Bounded concurrency (6) over the
~15-30 team fetches, Redis `espnroster:{sport}` (24h) + in-memory mirror,
fully defensive (a failing team is skipped → partial index, never throws;
grouped OR flat athletes[] shapes handled; non-numeric ids dropped). This is
now the PRIMARY resolver in resolveAthleteId/getPlayerGameLog; the v2 search
stays as a backstop on a roster miss. A unique roster hit wins (S59 doctrine)
— a missing name beats guessing another player's id.
2. getPlayerGameLog retries the gamelog with candidate seasons (current +
previous calendar year) ONLY when the first parse comes back empty
(filters-only), unlocking the dual-league athletes. The common path is
untouched.
MLB path (statsapi) unchanged; settlement/snapshot/frontend untouched.
Live probe: WNBA roster index = 206 players (Collier id 3917450 / Lynx team 8,
Brionna Jones id 3058895 present); NBA index = 544. Collier now resolves
end-to-end with 20 gamelog rows (was NOT FOUND); Brionna Jones likewise; all
previously-working players (A'ja Wilson, Ionescu, Clark, Stewart, Plum) still
resolve. NBA hyphen names (Gilgeous-Alexander) resolve via nameKey folding.
Tests: tests/unit/espnRosterIndex.test.js (9, fail-then-pass on base adapter).
Full backend suite 3183 green; web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
FREE ESPN news wire + championship-winner futures for the never-dark
offseason hub. Both graceful/empty, never fabricate a market value.
- newsService (mirrors injuryService): per-sport ESPN /news FEEDS, pure
parseNews → { sport, items:[{id,headline,description,published,type,
athlete?{name,key},team?,href}] }; athlete/team from categories[] only
(absent when not present). Cache 15m, injectable, offline-tested.
- oddsNormalizer.normalizeOutrights: NEW branch — outrights outcomes are
{name,price} with no point, so normalizeProps drops them; keeps them with
best-price-across-allowed-books per selection. + americanToDecimal.
- oddsService.FUTURES_KEYS: separate map (mlb/nba/wnba championship winner),
OUT of the daily SPORT_KEYS/snapshot budget.
- futuresService: getFutures(sport,deps) → { sport, updated_at, markets:
[{key,title,selections:[{name,price,prevPrice?,move?}]}] }. One outrights
call per 12h TTL (quota-disciplined), FUTURES_ENABLED gate. Price-move
(shortening/drifting/flat) mirrors computeLineDeltas SHAPE on odds not
line; prev prices persisted inside the futures:{sport} value (no new key).
linkNewsToMoves pure causal-tie helper.
- Routes /api/news/:sport + /api/futures/:sport (registered) + Next proxies.
- Tests: newsService, futuresService, oddsNormalizerOutrights (fail→pass,
no network). Full suite green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Extend /explore (ExploreHub) into the 365-day never-dark hub. Two new
self-hiding sections feed REAL always-available data into the offseason:
- NewsWire (components/vyndr/NewsWire.tsx): real ESPN headlines from
/api/news/:sport (newest-first, mono timestamps, type chips, player/team
links) + real injuries from /api/schedule/:sport/injuries (OUT/GTD chips,
token colors). Reuses the retired TerminalTemplates INJURY_WIRE layout but
never routes its sample constants. Self-hides when both feeds are empty.
- FuturesBoard (components/vyndr/FuturesBoard.tsx): real futures from
/api/futures/:sport — championship/win-total/award markets, mono tabular
prices + movement colored by the contract (shortening=green / drifting=amber
/ flat=dim, NEVER red; move shown ONLY when the backend supplies one).
Carries the honest "TRACKED · NOT GRADED" label — no fabricated grades on
futures. Self-hides when markets:[].
- ExploreHub is offseason-aware (via emptyState OFF_SEASON month check): the
hub LEADS with futures + wire when the board is dark, COMPLEMENTS the live
board in-season. Sport selector kept; each section self-hides independently.
- Testable pure helpers: lib/futuresMove.js (move→color, never red) +
lib/newsFormat.js (timeAgo mono-stamp, ESPN type labels).
Contracts consumed (Wave 2A owns the proxy/service files); code self-hides on
fetch failure if a proxy isn't present yet.
Tests: tests/unit/newsWire.test.js + futuresBoard.test.js (26 new). Full suite
259 suites / 3144 green; web build exit 0. vyndrParityQA stays green
(mono data, no glitch on data surfaces).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Unblocks the self-learning loop for basketball. Once an NBA/WNBA grade
exists (Wave 0), it now settles against the FREE ESPN per-game log
(espnStatsAdapter.getPlayerGameLog) — the same {found, last10:[{date,stat}]}
contract MLB settlement already consumes. accuracy:{sport} + by_tier
calibration + the Wave-3 TierRecord light up automatically.
- outcomeService/ledgerService: defaultGetPlayerStats routes nba/wnba to
espnStatsAdapter.getPlayerGameLog; MLB stays on mlbStatsAdapter.
- outcomeService: sport-aware statValue + a SEPARATE NBA_BOX_KEY/NBA_COMBO
map (S11 three-map-split kept — never merged with MLB_LOG_FIELD). Combos
(pts_reb_ast, reb_ast, stl_blk, …) sum components; a missing component
never fabricates a total.
- logRowOnDate: ESPN gamelog rows carry a FULL ISO timestamp (a late tip
rolls past UTC midnight), so basketball date-matches on UTC OR ET date;
MLB keeps exact YYYY-MM-DD compare. Outcome `date` is normalized to the
ET calendar day so the accuracy window filter + idempotency key behave
identically across sports.
- Final-honesty guard: never settle a basketball row whose ET date is
today (an in-progress partial box). MLB is final-only + settles same-day,
so the guard is scoped to basketball. The ledger path is already guarded
(.lt('game_date', today)) for all sports.
- opsWatch: nba/wnba added to SETTLEABLE_SPORTS; zeroSettleAlarm gates them
behind a real-finals probe (finalsBySport) so an offseason/off-day's
stale pendings never false-page "settled 0". snapshotScheduler counts
yesterday's ESPN state==='post' events and feeds the map; MLB unchanged.
- snapshotScheduler: boot announce per settleable sport
([settle:mlb] [settle:nba] [settle:wnba]). Thrown-error paging already
covers the new sports (settleAll* loop every sport).
Tests: tests/unit/nbaSettlement.test.js (16) — WNBA hit/miss/push, combo
pra, idempotent re-run, unplayed/today game does NOT settle, accuracy:wnba
+ byGrade + by_tier populate, ledger WNBA settle. opsWatch (+5) — finals
off-day no page, finals present DOES page. MLB suites unregressed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
featureCache.gameLogFeatures falls back to espnStatsAdapter.getPlayerGameLog
(free ESPN per-game logs) when the offline Python source returns null → NBA/WNBA
produce l5/l20 → props GRADE instead of refusing. +18 tests, proof test locks
the refuse→grade transition. MLB unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The Python nba_api service (gameLogService) is offline in prod, so
featureCache's non-MLB branch produced no l5/l20 averages →
projectionFor returned null → the ENTIRE NBA/WNBA slate refused
(insufficient_data). Only MLB actually graded.
Fix (free, no-auth, verified live):
- espnStatsAdapter.getPlayerGameLog(name, sport) — resolves name→ESPN
numeric athlete id via the v2 search (the v3 /search now returns
count:0; the v2 uid carries a:<id>, defaultLeagueSlug disambiguates
league), fetches the per-athlete gamelog, and parses per-game rows
keyed by VYNDR stat names (points/rebounds/assists/threes/steals/
blocks/turnovers + computed pra). Columns are indexed by the
response's own names[] array (NBA and WNBA orders DIFFER), never
positionally. Most-recent first, defensive (null on unrecognized
shape, never throws), cached (espngamelog:{sport}:{id} 4h + memory).
- featureCache.gameLogFeatures — falls back to the ESPN gamelog for
nba/wnba when the Python source returns null/empty, producing
l5/l10/l20 + rest_days + minutes_per_game via a new local
NBA_LOG_FIELD map + pure nbaGameLogFeatures (S11 three-map-split:
separate from MLB_LOG_FIELD).
Grade gates already whitelist all 8 NBA/WNBA stat types in both Node
paths (analyze.js + scan.js); no gate change needed.
Tests (hermetic, no network): espnGameLog.test.js (parser/resolver/
adapter) + featureCacheNba.test.js (the UNLOCK proof — empty features
refuse, ESPN-derived features grade). 3098 tests green; next build
exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Harvest ESPN athlete id + direct headshot href from the game-summary feeds
already called (independent of the flaky Python stats path) → reliable
NBA/WNBA/NFL/NHL headshots when in-season; soccer stays honest monogram
(no free id). MLB unchanged (zero extra I/O).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The NBA/WNBA espnId was captured only from espnStatsAdapter (the offline-
Python fallback), unreliable in prod. Add espnAthleteIndex — a pure,
defensive harvester that builds { nameKey -> {espnId, headshotHref} } from
the ESPN schedule->summary/boxscore/leaders/injuries/roster feeds the
pipeline already calls (free, bounded mapLimit, cached, MLB->{}).
snapshotService now fills any player the primary stats-resolve left without
an espnId from this index, and stores a DIRECT headshotHref as headshotUrl
on the enriched grade (the exact URL, never 404s on a constructed path).
Threaded headshotUrl through slateAdapter.buildPlayerStripsFromProps ->
GameCard -> StatStrip -> PlayerAvatar/getHeadshotUrl (direct href wins over
the constructed one). MLB's MLBAM path is untouched. Soccer resolves only
via a direct href; absent -> honest monogram (API_FOOTBALL_KEY remains the
reliable soccer path, unwired).
getGameSummary now also passes through ESPN `rosters` (pre-game lineups
carry id + headshot). Everything graceful: any miss -> absent -> monogram.
Tests: tests/unit/espnHeadshotIndex.test.js (11) — fixture->index, snapshot
merge fallback, direct-href-wins, soccer honest monogram, malformed/cyclic
parse never throws. Full suite 3080 green; web next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Real entity assets (headshots/logos/book wordmarks), record-by-grade-tier,
outlook mode, market-breadth, Parlay Lab, grade-shift timeline, /u house-mode
public record, pitcher arsenal (Savant), combat intelligence v1 (MMA), and
three trust-bug fixes. 253 suites / 3069 tests green, next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Reserved house handle (default 'vyndr', env HOUSE_HANDLE) resolves to the
PUBLIC model record — getModelAggregate() with no userId (user_id=NULL rows)
— WITHOUT a public_profiles row. It is the ONLY special case; every other
handle keeps the private-by-default, byte-identical-404 no-existence-leak
contract. The house profile is always public and never 404s (a fetch failure
degrades to an honest building state).
- src/routes/profiles.js: house short-circuit + sendHouseProfile (public
aggregate + by_tier + public settled entries), reserved before the publish
lookup so a user claim is shadowed.
- PublicProfile.tsx: house label 'VYNDR MODEL · PUBLIC RECORD' + hero/subtitle
off data.house; keeps the CLV-VERIFIED record hero + TierRecord calibration
+ recent settled reads (misses included).
- opengraph-image.tsx (1200x630): house-branded eyebrow/heading.
- portrait/route.tsx: new 1080x1350 share crop (real aggregate or tagline
fallback, never a fabricated number).
- Discoverability: 'VIEW AS PUBLIC PAGE ->' on the ledger MODEL header +
'VIEW PUBLIC RECORD ->' under the landing ModelRecord, both to /u/vyndr.
- tests/unit/houseProfile.test.js: house resolves to user_id=NULL aggregate
(no public_profiles row) + by_tier; unknown/unpublished user handles stay
byte-identical 404; page renders house label + TierRecord + portrait crop.
3019 tests green (3012 -> 3019); next build EXIT=0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Step 3 — OUTLOOK MODE. The game grid no longer dead-ends in a "NO SLATE" CTA.
When there are no live games (and it's not a network failure) it shows REAL,
always-available data: yesterday's PROVEN A-tier receipts (/api/ledger/model)
+ tomorrow's date-pinned ESPN schedule preview (free/cached). A network
fetchError stays a distinct ERROR state — never a fabricated outlook.
- lib/outlook.js (new, CommonJS, unit-tested): buildOutlook selection +
mapTomorrowPreview (upcoming-only, drops incomplete matchups, never invents).
- Slate.tsx: OutlookSurface replaces the empty-grid CTA (dateOffset 0 only).
- dashboard/page.tsx: DashboardOutlook replaces the "Today's games" NO-SLATE CTA.
Step 4 — MARKET-BREADTH / CONSENSUS vs MODEL. Makes the DeskShowcase
"consensus vs model" claim REAL. Consensus = median book line across a prop's
per-book rows; the model's position is model_value vs consensus, signed by the
graded side. <2 distinct books → null (never fabricate a consensus); a
non-numeric line is ignored, never coerced to 0.
- lib/marketBreadth.js (new, CommonJS, unit-tested): median/computeBreadth/
collectBreadth (strict null guards).
- components/vyndr/MarketBreadth.tsx (new): mono/tabular strip, colored by sign
via colorContract.edgeColor, self-hides when nothing has >=2 books.
- Slate.tsx renders it above the grid (joins books + snapshot model_value).
- slateAdapter.js exports gradeKey for the join.
- DeskShowcase.tsx: the consensus claim is now backed by the shipped feature.
Tests: tests/unit/outlook.test.js + tests/unit/marketBreadth.test.js (23 cases).
Full suite 2984 passing (245 suites); next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
TASK 1 — Parlay Lab (/parlay): a dedicated Correlation Builder with a
leg source INDEPENDENT of the live slate. Browses tonight's pre-graded
props from /api/snapshot/:sport (resolves players via /api/players/search),
adds legs through useParlay().addLeg (deduped by legKey), and renders the
PARLAY SLIP — combined grade, correlation, and payout read straight off
ParlayContext. Surfaces parlayService's correlation warning as the
CAUTION · CORRELATION FLAG, honors the tier leg-cap (free 2 / analyst 4 /
desk 6) and blurs the payout for free tier with the __goPaywall upsell.
Added /parlay to OPEN_ROUTES (free funnel, like scan/dashboard). Retired
the cleanly-dead ParlayTray.tsx (unmounted since Session 50). No new
proxies — reuses existing snapshot/search/parlay-grade endpoints.
TASK 2 — Live Grade-Shift timeline: web/src/lib/gradeShift.js (pure,
testable) builds a line/grade-movement timeline from already-emitted data
(intraday {t,line} history + revised_from_grade). Color law mirrors
ROW-GRAMMAR / StatStrip.LineSparkline: green = toward the graded side,
amber = against, dim = flat (never red). GradeShift.tsx renders it, shows
the original grade struck-through on a revision, and self-hides below 3
real points. Mounted in GradeResultCard (self-hides on the scan path,
which carries no captured history — honest, never fabricated).
Tests: tests/unit/gradeShift.test.js (15) + tests/unit/parlayLab.test.js
(13). Full suite 245 suites / 2989 tests green (baseline 243/2961).
Next build EXIT=0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ONE shared TierRecord component (lib/tierRecord.js + TierRecord.tsx) on
dashboard + /u + ledger — per-tier calibration (A+ X-Y, A X-Y, …), W-L always,
hit-% only at n>=20 per tier. by_tier flows through all endpoints untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Addition 2 (non-negotiable): the model's record must show PER GRADE TIER
(A+ went X-Y, A X-Y, …) everywhere the record appears. A blended % hides the
proof that higher grades win more — the tier calibration IS the credibility.
The backend already computed `by_tier` in getModelAggregate; this is display
propagation via ONE shared component (the class fix, not five one-offs).
- web/src/lib/tierRecord.js — testable CommonJS row-builder. W-L counts ALWAYS
(honest at any n); hit-% only when the upstream n≥20 gate passed (hit_pct !=
null), else "RECORD BUILDING · N settled". Order A+ A B C D F. A-tier is the
only edge (green) tier — no glow below A, red reserved for outcomes.
- web/src/components/vyndr/TierRecord.tsx — the ONE shared table. Presentational
(byTier) for /u + ledger; self-fetch (/api/ledger/model, sport-scoped) for the
dashboard. Fully self-hides until a tier has a settled read.
- Ledger swaps its inline TierCalibration for the shared component (single
source). /u PublicProfile renders it below the blended hero (by_tier added to
the aggregate type; it already flows through the route + proxy untouched).
Dashboard gains a compact per-tier surface.
- Endpoint/proxy audit: profiles.js + ledger.js return the full aggregate
(by_tier included); both Next proxies pass the body through — no threading
needed. No change to the gate or math in getModelAggregate.
Tests: tests/unit/tierRecord.test.js (row logic + tier order + edge contract +
source-assert all three surfaces render the shared component). ledgerService
test gains an A+-stands-alone bucketing case. Full suite green (243 suites /
2961 tests); web next build exits 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The merged suite exposed two self-defeating assertions in 2B's book test:
uppercasing the name before comparing to the key false-failed correct brand
names (DraftKings→DRAFTKINGS), and the lowercase-echo guard rejected bet365
whose official wordmark IS lowercase. bookInfo output was always correct.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Threads a REAL athlete id from the snapshot's per-player stats resolve (zero
new I/O) → enriched grade → grades:{sport} → slate strip → PlayerAvatar. Real
photo where an id resolves; team-colored monogram (never a gray silhouette,
never a broken image) where it can't. Ids are never fabricated.
Ingestion (Addition 1):
- espnStatsAdapter.getSeasonAverages now RETURNS the resolved ESPN athlete id
(was discarded) as espnId; non-numeric uid degrades to null.
- playerIntelService surfaces MLBAM playerId (MLB) / ESPN espnId (NBA/WNBA).
- snapshotService captures both per player and stores them on the enriched
grade beside archetype/team (null when unresolved → monogram path).
Thread → component:
- slateAdapter.buildPlayerStripsFromProps carries playerId/espnId onto each
strip; StatStrip → PlayerAvatar (accepts both ids; getHeadshotUrl routes by
sport: MLB→mlbstatic, NBA/WNBA→a.espncdn).
- Silhouette surfaces rewired to PlayerAvatar (branded monogram on null):
scan search dropdown (guarded MLBAM p.id) + tonight chips, SearchModal,
HotListPanel, GradeResultCard header. Scan grade card feeds the picked
MLBAM id through gradeAdapter.
- playerHeadshot pure URL logic extracted to CommonJS playerHeadshotUrl.js
(unit-testable; the .ts re-exports it). nfl/nhl added to ESPN_SPORT_PATH.
Tests: tests/unit/headshotThread.test.js (per-league URL + id thread + monogram
null path) + extended snapshotService/espnStatsAdapter suites. Full suite
241 suites / 2915 green; next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
MISSION 1 — real sportsbook wordmarks (kills lowercase "betmgm"):
- books.js: add the 6 missing ALLOWED_BOOKS keys (fanatics/bet365/
hardrockbet/betrivers/pointsbet/pinnacle) with real brand names +
colors — no live book falls to neutral gray. Add `slug` fields +
bookSlug()/hasBookSvg() + BUNDLED_BOOK_SVGS.
- Bundle 8 self-authored styled-text wordmark SVGs under
web/public/books/{slug}.svg (draftkings/fanduel/betmgm/caesars/
bet365/pinnacle/hardrockbet/betrivers). NOT copied trademarked logo
glyphs — the book's NAME in brand weight+color; official press-kit
art can drop into the same paths with zero code change.
- BookWordmark: render the local SVG when bundled, else the brand-color
styled-text fallback (never a broken image; never a lowercase key).
- Import BookWordmark into the ledger row (page.tsx:368) + the identical
public-profile row, replacing bare {row.book} text. vyndr/GameCard
line-grid book cell now proper-cases via bookInfo().name (keeps the
preferred-book green highlight).
MISSION 2 — team-logo coverage gaps:
- teamMeta.js: add ESPN-schedule ball-sport abbr aliases the feed emits
that fell to monograms — SA→SAS, NY→NYK, WSH→WAS, BRK→BKN (NBA),
CONN→CON (WNBA). Real-abbr-first lookup means MLB WSH (Nationals) +
WNBA NY (Liberty) still resolve directly; NY in MLB stays null.
Tests: new tests/unit/bookWordmark.test.js (all 10 ALLOWED_BOOKS resolve
to a real brand+non-gray color; 8 bundled SVGs exist; BookWordmark
SVG-first + no-lowercase-leak; ledger/profile import + use BookWordmark).
entityLayer.test.js extended for the new aliases. Full suite green
(239 suites / 2891 tests); next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
FIX 1 — Honest billing renewal render. VYNDR tiers are monthly, so a
`subscription_end` far in the future (the manually-seeded "RENEWS 6/9/2036"
founder row) is a comped/lifetime/seed value, not a renewal. New
web/src/lib/billingDisplay.js `classifyRenewal()` → date | none | lapsed |
unknown (strict Date.parse guard, MONTHLY_RENEWAL_MAX_DAYS=60). Profile page
renders the classified label for both the "Renews" stat and the
cancel-scheduled "Access ends" line — no raw far-future date. No DB row mutated.
FIX 2 — MLB namesake collision (James Wood → "Chicago Cubs"). searchPlayer now
collects ALL exact-nameKey matches instead of first-`.find`; a ≥2 collision
resolves ONLY via a confident teamHint (the prop's game participants, matched
against the cached /teams list with ESPN↔statsapi abbr reconciliation), else
refuses (null) — never guesses. The hint threads getPlayerStats →
resolvePlayerStats → snapshotService (built from each prop's home/away team).
Join invariant: a single-exact player whose team isn't in the hinted game has
its team DROPPED (null), so streaks/rosterlogs never tag a foreign team. Full
teamHint recovery shipped (not just the refuse fallback).
FIX 3 — DeskShowcase headline "A $1M terminal." → deadpan value-showing copy
"Every grade, every alt line, live." Prices ($44.99 / $34.99) unchanged.
Tests: billingDisplay.test.js (7), mlbNamesakeResolve.test.js (12,
disambiguation + join invariant + pure helpers), ds5PricingStates updated to
assert the new headline and no "$1M". Full suite green (237 suites / 2863
tests); web `next build` exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Part 6 #8 — Desk $44.99 is the hero tier. New DeskShowcase leads the
pricing page with the "$1M terminal · $44.99" story + real feature
visuals (alt-line ladder, quarter-Kelly, parlay φ, real-time feed) in a
balanced two-column layout (kills the dead right-half). Desk is the sole
highlighted tier / single primary CTA (color contract #9); Analyst is
secondary. Real prices: Free 5 scans, Analyst $14.99/$19.99, Desk
$34.99/$44.99. ClaimMeter + Stripe checkout wiring untouched.
Part 4 #7 — Ticker → punctuated stillness. The continuous marquee is
retired; the ticker now RESTS ≥4s on each ranked item and pulses only on
change. The EKG heartbeat is a static readout — the header's ONE idle
proof-of-life is the single SIGNAL-LIVE live-dot (the ticker dropped its
competing pulse). All durations tokenized (--motion-*, --ticker-hold);
prefers-reduced-motion kills the motion entirely.
Part 8 #20 — new EmptyState component modeled on the north-star 404
(scanlines + glitch wordmark + amber system voice + CTA hierarchy),
reused at the bare-red "Team not found", "Game not found", and the
ledger empties — one unified voice.
Part 5 — archetype glyph+chip propagated to STREAKS rows + ledger rows
(optional + self-hiding; absent beats fabricated); grade reveal already
carries ArchetypeBlend.
Tests: new tests/unit/ds5PricingStates.test.js (22) locks all four
workstreams; updated teamHubUI + vyndrDesignSystem for the new surfaces.
237 suites / 2864 tests green; next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
DESIGN-SPEC Parts 3 + 6 (audit #1, #13, #14). The founder's named #1 rebuild.
slateAdapter.js — the testable engine:
- selectTopGrades: rank tonight's grades by tier → confidence → edge so the
row varies on a real signal, not identical-weight noise (#13).
- buildHeroReceipts: yesterday's PROVEN A-tier settled HITS (misses excluded),
carrying the real result — the never-empty proof source (#1, Part 6).
- heroFallbackState: tonight wins, else receipts, else empty.
- pendingSummary: collapse an all-awaiting card's six "Grades post …" rows to
ONE line count (#14).
- topReadForCard: the single best live graded read to promote (#2).
GameCard.tsx — ONE bold hero per card (large mono/tabular grade + player, rest
demoted); all-awaiting cards render one "N props pending · grade ~X ET" line via
nextRunLabelET instead of repeated filler. Real team logos + team-colored accent
already lead the card (DS0) — preserved.
dashboard/page.tsx — Top grades tonight ranked via selectTopGrades (+ % CONF the
varying signal); when tonight is empty, fetch /api/ledger/model and fall back to
yesterday's PROVEN A-tier receipts (✓ HIT + actual + CLV) so first paint always
proves the model. Honest nextRunLabelET copy kept for the truly-empty case (QA.22).
Tests: tests/unit/ds2Dashboard.test.js (21) — pure-fn + source assertions,
fail-before / pass-after. Full suite 237 suites / 2863 tests green (+21).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The OAuth-only 'sb-token' localStorage key was read by profile, slip,
dashboard (recent-scans), settings, and tracker for their authenticated
fetches. Email/password users never had that key, so those fetches sent
no Authorization header and silently returned nothing.
- web/src/lib/authToken.js — currentAccessToken() reads the REAL Supabase
session (sb-<ref>-auth-token, v2 top-level or v1 currentSession), legacy
fallback. CommonJS so Jest can unit-test it (5 tests).
- Swept all 5 pages to the helper (scan already session-first from DS1).
- lib/api.ts (0 callers) + ParlayTray (unmounted) left as dead code.
236 suites / 2842 tests green, next build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Both sessions fixed the negative-edge-green bug; kept DS3's centralized
colorContract.edgeColor helper (removed DS4's local const), updated DS4's
source-assertion test to match. Same semantics, one enforcer.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signal-green #00D4A0 now means exactly ONE thing (edge/active/A-tier/CTA),
locked by tests that fail on violation. Part 1 of DESIGN-SPEC v2.
- web/src/lib/colorContract.js — pure CommonJS helpers: edgeColor(value)
colors edge/CLV/delta by SIGN (neg=var(--miss), pos=var(--g-a), 0=neutral);
gradeTierColor() (A/A+ green, B blue, C amber, D/F red, in lockstep with
vyndrTokens.gradeColor); gradeGlows() (A/A+ only); deltaE()/isSignalGreen()
CIE76 gate so no archetype hue dilutes the signal.
- GradeResultCard: edge confidence-strip + EDGE row route through edgeColor
(a -33.3% edge was rendering GREEN — audit #3); grade-hero glow gated to
A/A+ via gradeGlows (a glowing C devalued the cue); VYNDR INTELLIGENCE
panel de-flooded (neutral border, Form/Rest neutral not green — #15).
- LiveHeroProp: negative edge now muted red, not neutral (sign completeness).
- Archetype dedup off signal-green (both archetypes.js + archetypeService.js,
kept matched): DUAL THREAT/MOTOR #00D4A0, MIRROR #34D399, ARTILLERY/RANGE/
GHOST/BLADE #2DD4BF, BRUSH #3DDC84 shifted to distinct non-green hues (all
ΔE>=44 from #00D4A0). Within-sport uniqueness preserved.
- tests/unit/colorContract.test.js — 21 tests: helper units + source-grep
violation locks + archetype-green-dedup. QA.20-22 kept green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Semantics preserved (step-up amber / step-down green); DS4 moved the
coloring into a sev object, so the source-lock regex is updated to match.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
P0 billboards (the timeline is customer #1). Pixel-level craft: one bold hero
among muted context, real entities from DS0.
- STREAKS row: rebuilt from a log line into a Bloomberg alert. Streak LENGTH is
now the mono/tabular HERO (38px, the largest figure in the row); real
PlayerAvatar identity; muted "built vs [opponents]" lens; ONE severity accent
(step-up amber / step-down green); grade badge tier-gated (the READ is paid).
- Grade reveal: edge is now sign-colored — negative edge uses var(--miss),
never green (color contract #3). Fixed in BOTH the confidence strip and the
MODEL/LINE/EDGE row; that row is now mono + tabular. Letter stays the hero.
- CLV reframe: new pure lib/clvDisplay.js (clvMode flat/spread/none). A near-
flat distribution (73/74) now renders a confident VOICE line — "CLV flat — we
grade the outcome, not the close" — instead of a broken-looking histogram;
bars show only on real spread.
- /u profile: hit% is the bold record hero (56px mono tabular) + beat-close
secondary; honest CLV-VERIFIED badge (only when closing value is tracked);
real PlayerAvatar identity on cards; OG image elevated to carry the real
record at 1200x630 social crop (graceful tagline fallback).
Tests: +23 (tests/unit/ds4Billboards.test.js). Full suite green (2732).
web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Fixes the three DESIGN-SPEC Part 4 + #17 audit findings.
1. React #418 hydration mismatch (landing → dashboard entry). The
`maybeSignedIn` value was computed in a useState INITIALIZER that reads
localStorage during render: server (no window) → false → emits the
marketing tree; a signed-in visitor's first CLIENT render → true → emits
the loading placeholder. Whole-subtree server/client mismatch → React
discarded and re-rendered the page. Deferred behind a mounted flag so the
first client render matches the server; the stored-session check flips
post-mount. SSR HTML is no longer discarded.
2. Loading walls → skeletons. New tokenized Skeleton primitive
(.vyndr-skeleton, reduced-motion-safe via the global rule). Swapped into
every text-wall loader: dashboard slate load ("Loading the slate…"), /desk
("Assembling the pack…"), /ledger ("Loading…"), scan ("Loading the model…"),
and the landing redirect placeholder. No bare text loader remains.
3. scan→ledger persistence. Root cause: the scan page read its bearer token
from localStorage['sb-token'] — a key written ONLY by the OAuth callback —
so email/password users posted /api/scan anonymously and the ledger write
(gated on an authed user) was silently skipped. Now uses the authoritative
session.access_token (matching the ledger read path). Extracted the row
builder to web/src/lib/ledgerRow.js (shared, testable).
Tests: +17 (scanLedgerPersistence write→mine round-trip + scope + idempotency;
ds1SpeedTrust hydration/skeleton/persistence source invariants). Full suite
233 suites / 2793 green; web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Founder's #1 priority. A single cached asset+rendering system, swapped into
every surface, so entities stop being flat gray strings (DESIGN-SPEC Part 2).
- web/src/lib/teamMeta.js: static registry for ALL 4 sports — 30 MLB, 30 NBA,
13 WNBA teams + 48 World Cup national teams, each with real colors + the
ESPN logo/flag CDN abbr. resolveTeam (abbr/full-name/nickname/alias),
teamLogoUrl (statsapi->ESPN mapping: AZ->ari, CWS->chw; soccer via the
countries/ flag CDN), accentColor (picks the VISIBLE color of the pair so a
#000000 primary never renders an invisible accent on #06060B). Colors +
abbrs sourced once from ESPN's team API — stable public facts, zero-latency
static data, no paid dependency.
- TeamLogo: real ESPN-CDN logo with a team-colored MONOGRAM fallback (never a
gray box / bare abbr). PlayerAvatar: real headshot with a team-colored
monogram fallback (kills the gray silhouette). BookWordmark: brand-color
wordmark, proper casing (DraftKings, not 'draftkings').
- Swapped into the class-level shared components so it propagates to ALL
surfaces: GameCard (team logos + team-colored accent edge), StatStrip
(player identity avatar), StreaksPanel (P0 billboard avatars), TeamHub
header (the team's real crest leads its hub). Barrel-exported.
ZERO OUT-OF-POCKET: ESPN logo/flag CDN + league headshot CDNs, all verified
200 image/png across MLB/NBA/WNBA/soccer.
ACCEPTANCE (SSR render proof): /entity-demo harness server-rendered the exact
real asset URLs across all 4 sports — mlb/500/nyy.png, mlb/500/chw.png (White
Sox, correct ESPN abbr), nba/500/lal.png, wnba/500/ny.png, countries/500/
usa.png + bra/eng/arg/jpn flags, real mlbstatic/nba headshots (Judge 592450,
LeBron 2544), DraftKings/FanDuel wordmarks. Each URL curl-verified 200
image/png. Harness removed post-proof (not a product surface). Pixel
screenshot blocked by WSL2<->Windows-Chrome localhost networking, not code.
2757 -> 2776 tests (+13 entityLayer, +6 boot resilience), web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
ROOT CAUSE: the Dockerfile copied src/poller/scripts/supabase but NOT
content/. mediaEngine.js read content/stark-lines.json with an unguarded
module-load readFileSync; ENOENT in the image threw at require time, and
via app.js → routes/desk → deskService → mediaEngine that crashed the
ENTIRE API at boot. The Coolify healthcheck rolled back to the last
healthy image (4d2b27d), so every deploy since 219167e silently served a
14-hour-old build — S11 live tracking, S6 API code, the settlement boot
line, and SNAPSHOT_EXPECTED_INTERVAL were all merged but NOT running.
FIX (one train):
1. Dockerfile COPYs content/ into the runner image.
2. mediaEngine: stark-lines.json is OPTIONAL (garnish, never load-bearing)
— loadStark() try/catch → {} → posts render without the Stark kicker,
never a crash. Belt AND suspenders with #1.
3. src/preflight.js (§A4): boot prints '[preflight] OK' or 'DEGRADED'
naming exactly what content/env is missing — before the healthcheck
can fail silently. Run first in server.js.
4. Full fragility sweep: mediaEngine was the ONLY unguarded module-load
file read; coachSignals (config/coaches.json) was already lazy +
try/catch + copied. No others.
Verified: requiring app.js + deskService + mediaEngine with
stark-lines.json ABSENT now boots clean (reproduced the exact prod
failure). 2757 -> 2763 tests (tests/unit/bootResilience.test.js).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The session-opening doc: 3 trains shipped (2757 tests), live infra map,
pending Coolify env vars + entity placeholders, honest open items.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MLB statsapi + WNBA ESPN live boxscores -> per-player current values
(live:{sport}:{date} TTL 90s, /api/live/:sport + Next proxy). Pure
propState math (HIT / ON PACE / NEEDS N / HOLDS / LINE PASSED — never
red in-progress), attachLiveProgress strip join on nameKey+statType,
proximity-to-hit slate float, StatStrip LiveTracker in the ROW-GRAMMAR
outcome slot (spec amended + lock test updated). Grades never change
in-game — tracking, labeled as such. 2698 -> 2757 tests, web build 0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-12 00:45:29 -04:00
761 changed files with 92405 additions and 2776 deletions
- **Integration:** 3 new blueprints registered in app.py (coaching_bp, redistribution_bp, unconventional_bp), evolution + odds_scanner extended with new endpoints
---
## Session — Under-querying vs out of data (2026-08-05)
**Shipped**
- `scripts/backfill-context.js` — platoon splits backfilled to all 380 settled
hitters (was 298; ingest had only ever seen tonight's lineups).
- `scripts/reconstruct-game-environment.js` — joins the ledger's game slug to
statsapi, writes `game_context` on the LEDGER's key, pulls actual archived
Open-Meteo weather. 96/101 settled games now carry real weather; park
dimensions 15 -> 30 venues.
- `src/services/model/parkWeather.js` (+ tests) — park geometry + air read onto
HIT TYPE, not P(hit). Wind refused (no park orientation).
'Which fair-probability ruler produced fair_prob_lock. NEVER pool edge or CLV across differing values — the denominator changed. v1_first_book = first admitted book (incumbent); v2_consensus = median across >=2 reference books at the same line.';
| 23 | Article media / share (S3) | YES | PARTIAL | NO | PARTIAL | YES | ❌ | `Intelligence.dc.html` ARTICLE MEDIA; `ShareCard.tsx`**0 real importers = DEAD**; OG `opengraph-image.tsx` IS live. In-article archetype figures not built |
| 24 | Calibration / edge board | YES | YES | NO | NO | YES✓ | ❌ | ✓HONESTY PASS: placeholder-edge% `MobileEdgeBoard` REMOVED from the Slate (phones show real cards). Component kept as dead code until a real edge feed exists |
| 25 | System / Intelligence terminal | YES | YES | **NO** | PARTIAL | YES | ❌ | `System.dc.html`/`Intelligence.dc.html`; `/terminal`→redirect to `/dashboard`; `/intelligence` REAL but **orphan (0 nav links)**; `/system` no page (prod 404) |
| 26 | Offseason hub (S-2) | YES | PARTIAL | NO | NO | — | ❌ | `Offseason.dc.html` full spec; **no `/offseason` page** (prod 404); logic only inline in `FuturesBoard`/`NewsWire` on Explore |
Prod endpoints confirmed LIVE + real: `/api/snapshot/{mlb,wnba}`, `/api/accuracy` (n=763),
## MODEL MATRIX — champion serves; every challenger is ledger-only
| Component | EXISTS | PROVEN | PROMOTED (serving) | USED (a surface reads it) | Evidence |
|---|---|---|---|---|---|
| Champion (engine1) | YES | PARTIAL→**promising** | YES | YES | `analyzeViaEngine1.js`; `enriched`→`snapshot`/`grades`. SKEW AUDIT 2026-07-29: on takeable MLB overs (n=62) champion p_win→CLV **partial r=0.375, SIG p≈0.003**; SURVIVES the mechanical baseline (no-edge CLV +1.5pt n=20 vs high-edge +8.6pt n=37 → **+7.1pt marginal**). De-vig clean (same-book pairing); close well-defined (DK/MGM r=0.92). PROMISING, NOT confirmed (thin n; lock-staleness check BLOCKED; 1 sig result among many) |
| arch-v1 | YES | NO | NO | NO | `challengerProjection.js`; rides `withChallenger`→**ledger only** (`:693`). "measured, never served" (`ledgerService.js:251`) |
| contact-v1 | YES | NO | NO | NO | `contactChallenger.js`; ledger col `p_win_contact` only |
| proj-v1 / v1.1 | YES | **NO (tested 2026-07-29)** | NO | NO | `projectionChallenger.js` (MLB-batting only). PROOF ORDER verdict: **NOT PROVEN** on n=45 takeable MLB overs — edge-CLV partial-r (controlling price) = 0.245 (n.s.); ~half the raw signal is the shared −fair_prob_lock term (mechanical); and the CHAMPION out-predicts it (champ partial-CLV 0.380 sig, champ-edge→hit 0.25 vs proj 0.12). Ledger-only |
| Champion alt-ladder | YES | NO | YES | Desk-only | `analyzeViaEngine1.js:486`; real re-grades ±1 line; `gradeAdapter.js:100` maps to card **Desk-gated** |
| proj-v1 `proj_ladder` | YES | NO | NO | NO | `distribution.js:100`; ledger-only, reaches no card |
| Price gate / EV | YES | PARTIAL | PARTIAL | PARTIAL | fields on `enriched`; hero gates on `isTakeable`+`ev_pct` (`heroPropService.js:78`). **But prod grades show ev_pct/p_win/model_odds/value/takeable = NULL** — built, served-schema, not producing |
### Ladder question (Phase 2.6) — VERIFIED: proj-v1.1 DOES compute rungs above the line
`projection/distribution.js:100``ladder()` computes `P(stat ≥ k)` for a fixed rung set
`k = 1..LADDER_MAX` (default 4), **independent of the listed line** — so for a line of 1.5
(tradedRung=2) it emits rungs 1 (below), 2 (at), **3 and 4 (ABOVE)**, each a real
— sole cap is `LADDER_MAX=4`. **BUT `proj_ladder` is ledger-only and reaches no user
surface.** So "ladder up from the listed line to find value" is *computed and never wired*.
## HONEST-STATE SUMMARY
### 3.7 — What a PAYING USER sees right now that is not true
1.**`/compare` — fabricated grades, nav-linked + public.** Hardcoded Jokić A+ / Wembanyama A + fake "VYNDR VERDICT" (`compare/page.tsx:10-62`). No data behind it. **The single worst live lie.**
2.**FAQ — phantom processor.** "We use NexaPay" (`FAQ.tsx:28`); every legal/pricing page says Stripe. Also founder price inconsistency ($24.99 vs $19.99).
3.**FAQ + Features — Brier/CLV over-claim.** "Brier score and CLV… published / from day one · Public accuracy by tier" (`FAQ.tsx:498`, `Features.tsx:323`). No Brier surfaced anywhere; CLV held. `CANNOT DETERMINE` a live Brier surface — because none exists.
4.**Mobile edge board — placeholder edge%.**`MobileEdgeBoard.tsx:44` renders a miscalibrated edge feed, masking >40% to "—". The numbers ≤40% still come from a placeholder pipeline. Live on mobile Slate.
5.**Price triplet / EV markers advertised-and-absent.** Grade schema carries `ev_pct`/`p_win`/`model_odds`/`value`/`takeable`; all **NULL on live grades** — the "model price" leg and VALUE marker don't render though the design promises them.
Not inflated (verified honest): grades are B/C only with A/D/F below n≥20 → `pct:null` everywhere; `AccuracyBadge`/`ModelRecord`/`TierRecord`/ledger all honor the n≥20 gate and self-hide. Hit-rate (59%, n=763) shows **without ROI/CLV** (CLV frequently null) — thin, not false.
### 3.8 — The graveyard (built, no user can reach it)
1.**`BookComparison.tsx`** — 0 importers. Backend `/api/books` is live + honest; nothing renders it. (S2 headliner.)
2.**`ShareCard.tsx`** — 0 real importers (S3 share cards).
3.**`components/GameCard.tsx` (legacy)** — type-only import; superseded by `vyndr/GameCard`.
4.**proj-v1 `proj_ladder`** — the above-line probability ladder; computed, ledger-only, never served.
5.**arch-v1 + contact-v1** — challengers, ledger-only, never served ("measured, never served").
6.**`TerminalTemplates.tsx`** — all SAMPLE data; `/terminal` redirects → effectively unrouted.
7.**`DemoScan.tsx`** — defined, never rendered.
8.**`/intelligence`** — REAL signals feed, but 0 nav links (orphan; reachable only by typing the URL).
9.**`/soccer`, `/marketplace`** — REAL pages, 0 nav links (orphans).
Tier-record · Pricing · Public profile `/u` · Player profile · Team hub · Explore ·
WIRE/Ticker · Streaks/Hot list · Live tracking.
## NEAREST TO DONE (one column from YES) — the shortlist a single order could finish
| Surface | The one gap | Finish move |
|---|---|---|
| **Book comparison (S2)** | WIRED: NO | Route `BookComparison.tsx` onto the card from the live `/api/books` store (crown stays off — measured flat). *Backend already shipped.* |
| **Parlay lab** | DESIGNED: PARTIAL | Accept the System "PARLAY BUILDER" spec as the lab spec (functionally live) — a doc call, not a build. |
| **Compare (H2H)** | HONEST: NO | Replace the hardcoded SAMPLE with a real two-player fetch, or pull it from nav until real. |
| **Price triplet** | HONEST: PARTIAL | Make the price-aware EV layer actually produce `p_win`/`ev_pct`/`model_odds` on served grades (built, not firing). |
Removed/hid every KNOWN live fabrication. REMOVE/HIDE only — no grade, snapshot,
scorer, pipeline, or real feature touched. Updated HONEST cells:
| Item | Was | Now |
|---|---|---|
| /compare (row 20) | HONEST: NO — hardcoded Jokić A+/Wembanyama A + fake VERDICT, nav-linked | Honest **in-development** state; removed from Nav + BottomTabBar. Real two-player build **pulled, awaiting real build**. |
| Pricing (founder copy) | $34.99 desk / struck $19.99 / FAQ $24.99 — wrong | Founder **Desk $44.99** (matches lib/checkout.js), **Analyst $14.99**; struck "regular" numbers removed; DeskShowcase $34.99→$44.99. First-100 counter is REAL (ClaimMeter→Stripe). No "first 50" desk claim (no such counter). |
| FAQ processor | "NexaPay" | **Stripe** (verified live: Next→Express→checkout.stripe.com). `nexapay.ts` + its webhook route were **PURGED 2026-07-27** (NexaPay Purge order) — cross-project contamination, never a real VYNDR path. Provider-side env/keys + the orphaned `user_profiles.nexapay_customer_id` column flagged for Kev. |
| FAQ + Features "Brier/CLV published from day one" | INFLATED (not surfaced) | Removed. Returns when Brier/CLV are actually surfaced. Backend Brier compute untouched. |
| Calibration / edge board (row 24) | HONEST: NO — MobileEdgeBoard placeholder edge% (masked >40%) | **Removed** from the Slate; phones show the real game cards. Component kept as dead code (hidden, not deleted) until a real edge feed exists. |
| Price triplet (row 19) | HONEST: PARTIAL — null model/EV rendered "MODEL READ WITHHELD · poisoned" (false quarantine) | New **NO_MODEL** honest-absent state: MODEL "—" / "NOT PRICED", no verdict. Fixes grade card + LiveHeroProp. EV layer still doesn't *produce* values (separate build). |
## KEPT ON THE BOARD (real work to finish — NOT cut)
- **Article media / S3 (row 23)** — real feature + free SEO/distribution. Finish, don't delete. Only the false "Brier/CLV" claim about it was corrected.
## FUTURE MODEL INPUT (logged only — not built this order)
- **News / line-movement signal** — injuries, scratches, lineups, weather move props before books reprice. Wire later as a model input. This is **ADDITIVE** to the media surfaces, not a replacement for them.
## KNOWN HONESTY GAPS (not fixed this order — logged, not fabrication)
- **Hit rate 59% (n=763) shown without ROI/CLV** — thin, not false. ROI/CLV surfacing is a later build.
- **"0 pushes = mis-scoring" — RETIRED 2026-07-29 as a false alarm** (premise re-verified, report-only).
The displayed hit/miss denominators are NOT corrupted by a hidden push bug: the feed is still 100%
173 lock_lines), all 992 settled actuals are integers, and the smallest actual-vs-line gap in the
whole ledger is 0.5. Expected pushes = exactly 0. See the verdict block below.
- **CLV instrument REPAIRED 2026-07-28** (commit 6552281). Was: 59 usable closing_prob. Now: **406** (MLB 248, WNBA 158) — the collapse was `attachClosingProb`'s `.limit(50000)`/no-ORDER-BY read + write-once `market_unavailable`, NOT capture (95% per-prop coverage) or the join (0 key mismatches). **CLV finding, straight: MLB unders lag the close (mean −9.1 prob-pts, 74% lose); MLB overs +2.0; WNBA flat.** → the +4.57% MLB-C and over/under asymmetry are substantially stale-line artifacts. This UNBLOCKS the proof order (proj-v1.1), which gates promoting p_win/ev to served grades.
**Honest state after this order: "no KNOWN live fabrications" — not "provably none."** The audit was thorough (repo + prod), but absence of a claim of falsehood is not a proof of universal truth.
**N-gate PASSED: overlap = 45** (settled ∩ proj-v1.1 ∩ MLB over ∩ CLV close ∩ fair_prob_lock ∩ takeable −160..+200). The CLV repair is what made n≥30 reachable. Edge basis = `proj_p_over_line − proj_book_implied` (de-vigged fair, VERIFIED not raw book).
**VERDICT: NOT PROVEN.** proj-v1.1's takeable-edge does NOT beat the champion or clearly beat the close on MLB overs.
- Phase 1 buckets (descriptive; only the neg bucket clears ≥15): 6%+ edge (n=11) shows hit 70% / fair-ROI +0.57 / CLV +12.95pts — but even negative-edge rows show +2.7pt CLV (whole over-side is elevated = the stale-high concern).
- Phase 2 partial correlation (the real test): raw r(edge,CLV)=0.455 → **partial r(edge,CLV | price)=0.245, n.s.** at n=45 (t≈1.64, p≈0.11). ~half the raw signal is the shared −fair_prob_lock term (mechanical). proj predicts the close itself only weakly (partial r(proj,close|lock)=0.281, n.s.).
- Shown, not judged: MLB unders (n=20) CLV −9.34pts, r(edge,CLV)=−0.03 (contaminated, zero signal); WNBA — proj-v1.1 does not run (MLB-batting only), N/A.
**Conditional dependency (moot):** the verdict was to be conditional on the under-capture audit clearing the over-side. It's moot — proj-v1.1 fails Phase 2.6 (loses to the champion) BEFORE the audit applies, so NOT PROVEN regardless. The stale-high concern is corroborated (whole-over-side baseline +CLV; half of proj's signal mechanical).
**Notable:** the one statistically-defensible edge signal here is the **CHAMPION's** p_win predicting CLV on takeable MLB overs (partial 0.380, p≈0.01) — NOT proj-v1.1. That champion signal is itself still audit-gated (if MLB overs are whole-side stale-high, even it is suspect).
**What accrues a re-test:** ~20-25 more settled takeable MLB-over rows (to test proj's residual ~0.245 partial against zero), AND proj-v1.1 must demonstrate it beats the champion — which it currently does not. Promotion stays HELD.
**VERDICT: SURVIVES BASELINE.** The marginal (+7.1) is ~5× the mechanical floor (+1.5) and the price-controlled partial correlation stays significant. The skew is essentially ONE-SIDED (unders lag −7.0 overall; overs carry only a small +1.5 floor, NOT the +7 a symmetric two-sided over-skew would show). De-vig is CLEAN (`analyzeViaEngine1.js:539` pairs over+under from the same book/fetch — no fresh/stale pairing). Close is well-defined (draftkings vs betmgm over-prob r=0.921, n=151).
**→ GREENLIGHTS building the takeable-edge grade ON THE CHAMPION (engine1 p_win), NOT proj-v1.1** (which lost the proof). This is the project's first edge signal to survive an adversarial audit.
**FLAGGED — promising, NOT confirmed:**
- Thin n (62 overs / 37 high-edge / 20 baseline).
- Phase 2 lock-staleness check is **BLOCKED**: multi-book lines AT LOCK are not retained (`bookprices` is Redis current-only), so we cannot fully rule out that part of the baseline is lock-time staleness. The de-vig being clean + the small baseline make a large hidden skew unlikely, but it's not excluded.
- Sharp (pinnacle) reference covers only **8** props (sharp-CLV +3.17pt, directional hint only).
- This is one significant result among many computed this session — do not overstate.
**What would strengthen it:** retain multi-book at lock (enables the sharp/consensus lock-staleness check), and accrue more settled takeable MLB-over rows. Held: no promotion, no served p_win/ev, no capture fix — diagnosis only.
**UPDATE 2026-07-29 (commit c7067c8): the lock-multi-book gap is now CLOSED.** New `lock_lines` table (migration 033, applied+tracked) persists each graded prop's per-book lines at the lock moment (`lockLineCapture` in `snapshotService`, fenced RLS-service-role-only, grade byte-identical proven). This unblocks the staleness audit for FUTURE rows — it does NOT retroactively fix the existing 62. Confirmation still needs weeks of accrued lock+close+outcome. Populates from the next snapshot tick.
The landing/hero (matrix row 1) selection was silently broken: it ranked on `ev_pct`, which is NULL on served grades, and **`Number(null) === 0`** made every prop tie at EV 0 → the "top read" was the FIRST takeable A/B prop in cache order — **arbitrary, dressed as ranked** (prod served Kelsey Mitchell, the #6 read by p_win). FIXED: rank by the **champion's p_win** (the only promising edge signal) among A/B **takeable-priced** reads (`isTakeable`−160..+200, same band as the proof/audit); strict-null guard; takeable filter excludes chalk; **no backfill** → honest empty state when nothing qualifies. p_win is ranking-only (never exposed; the route strips it). Display-only — reads caches, writes to nothing. No proven-edge/+EV/best-bet claim, no CLV/ROI/edge number. This makes the champion's p_win a real (display) consumer for the first time. Fingerprint VERIFIED: hero is the max-p_win read across sports (WNBA A), not old code's first-in-order MLB pick (Schanuel −135); untakeable chalk excluded. Visual auth-gated → data fingerprint.
`player`+`stat`, a finite `line`, uppercases `sport` for `SportPill`, and drops unrenderable rows.
**One shared ranking definition:** new `src/utils/gradeRanking.js`; `heroPropService` now imports
`takeablePWin` (was an inline copy, behaviour unchanged), the selector imports `rankGrades`, and the
web mirror is cross-checked by test. Board is grade-first ("top GRADES"), hero is p_win-first ("top
read") — they differ by design and agree within the leading tier.
**Honest limit:** the Next proxy forwards no Authorization and caches under a shared key, so via the
dashboard every viewer receives the free-tier payload — correct order, no paid values. That is the safe
default (forwarding auth into a shared cache is how paid payloads leak); per-tier delivery through the
proxy needs a tier-keyed cache and is NOT done here.
**Verified:** 311 suites / 3882 tests green (18 new, leak test on POPULATED p_win), web build exit 0.
**Post-deploy fingerprint (`72a14dc`) — the boundary proven in production:** 404→200 captured across the
deploy boundary; the anonymous live order is *1. Brionna Jones (edge 29.4) · 2. Rhyne Howard (edge 42.9)*
— an edge-only sort would lead with Howard (and the local induction over the stripped snapshot did), so
**server-side p_win ordered it** (Brionna .90 @−106 > Howard .745 @−120) while the payload carries
**PAID FIELDS: NONE**. A bogus bearer token also yields no paid fields (fail-closed proven in prod).
The board's proxy path now returns 10 props, so TOP GRADES **renders** instead of falling back to empty.
The rendered board is client-side → tagged for the Chrome audit, not faked.
---
# MODEL ARCHITECTURE RECOVERY MAP — 2026-07-30 (report-only) → `specs/model-architecture-recovery-map.md`
**The live grade uses 0 of the 3 specced layers.** Every audited metric (calibration, CLV, skew audit,
takeable floor, p_win→CLV r=0.375) is measured on the SHADOW model. Those findings stand — the shadow
model served every real grade — but none are evidence about the specced engine, which has never been
measured.
| Specced component | State |
|---|---|
| Layer 1 Similarity (`python/utils/similarity.py`, 101 ln) | BUILT · NOT WIRED · **NOT DEPLOYED** |
| Layer 2 Bayesian (`python/utils/bayesian.py`, 320 ln) | BUILT · NOT WIRED · **NOT DEPLOYED** — and its "sport-agnostic math / per-sport parameters" claim is TRUE of the built code |
| Layer 3 grade scale (`grade_thresholds.json`) | BUILT · **WIRED BACKWARDS** — the table maps PROBABILITY→GRADE; the live JS reads it in reverse to manufacture `confidence` from an already-chosen letter |
| Live champion (`engine1` + `probabilityEstimator`) | BUILT · WIRED — but the letter is a factor-index with **zero probability input**, and `p_win` is computed separately and never feeds it |
**The Python engine is not in the deploy image at all** (no python/pip in `Dockerfile`; `app.js` only
health-checks it). **The sport boundary is NOT clean on the live path** — adding a sport is a ~10-file
core edit with four documented silent-failure modes, so "make a sport a module" is itself a prerequisite
build. **Per-sport records DO exist** (`sports.mlb` n=526/62% vs pooled `overall` n=937/58%, each with
its own n≥20 gate) — the rule to enforce is that a new sport renders `sports.{sport}`, never `overall`.
**Park×weather is a CHALLENGER, not the champion** ("measured, never served"); xwOBA and bullpen-leash
are absent entirely.
Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting
echo"off-box push FAILED (on-box dump is durable, but OFF-BOX IS REQUIRED)"
echo"OFFBOX_OK=0"
notify "VYNDR OFF-BOX PUSH FAILED""urgent""Dump ${STAMP} (${SIZE} bytes) is on the persistent volume but did NOT reach the Storage Box. The database has no off-box copy tonight."
fi
else
echo"off-box push DEFERRED (BACKUP_OFFBOX!=1 or remote/key unset) — on-box dump is durable at ${DUMP}"
measurement:'EXACT ANALYTIC ABLATION of the champion, on the REPAIRED settled set. No refit, no re-fetch, no lookahead.',
limits:{
base_recency_not_separable:'stored features hold AVERAGES, not frequencies-over-line; reported as one block',
clamped_rows_excluded:clamped,
},
total_rows:rows.length,
per_stat_ablation:perStat,
residual_signal_test:residual,
archetype_test:archetypeTest,
multiple_comparisons_note:'The residual scan runs 14 unused features x 5 stats = 70 tests at alpha .05, so ~3-4 CI-excludes-zero results are EXPECTED BY CHANCE. Treat a single hit as noise; only a feature repeating across independent stats is evidence.',
mechanism:'Extra bases need BOTH conditions: hit hard AND hit in the air. A 105-mph ground ball is an out; a 25-degree popup is an out. Neither factor alone predicts bases, which is precisely why each may fail solo and the product may not.',
mechanism:'A hitter only realises his contact quality against a pitcher who permits contact quality. Elite suppression should attenuate a power bat; a contact-permitting arm should amplify it. The effect is conditional by construction.',
mechanism:'ARCHETYPE-CONDITIONAL. Barrels convert to extra bases for hitters whose lane is power; for a speed/contact profile the same barrel rate is a rarer event on a swing built for something else. This is Discipline 2 stated as a testable interaction. NOTE: it is currently UNTESTABLE — statcast rows carry no archetype label, and the barrel-relative proxy is an exact linear function of barrel_pct, so controlling for both components is rank-deficient. It needs a real archetype classification joined in.',
mechanism:'DEFENCE. A ball in play becomes a hit or an out partly by who is standing behind the pitcher. This should matter MOST for hitters whose value is contact that stays in the park, and LEAST for power hitters whose barrels clear the defence entirely — so a DEAD result for BOMBER is not a failure, it is the differential the theory predicts.',
mechanism:'DEFENCE x BATTED-BALL PROFILE. A low-launch (ground-ball) hitter puts the ball where fielders range; a high-launch hitter does not. Launch angle stands in for the profile, so defence should condition the ground-ball hitter far more.',
mechanism:'PITCHER BATTED-BALL TYPE. A ground-ball arm takes the air away, and a hitter whose value lives in the air needs the air. An air hitter against a sinkerballer and a ground-ball hitter against a fly-ball arm are both mismatches that neither factor states alone.',
mechanism:'ARSENAL MATCHUP. Barrel rate is far more a fastball skill than a breaking-ball skill, so a power bat facing a breaking-heavy arm should convert less of it. The pitch mix is already ingested, so this costs nothing to test.',
mechanism:'Strikeout risk compounds multiplicatively (log5 is exactly this shape). A high-K bat against a high-K arm loses plate appearances to strikeouts, and a PA lost is a base opportunity that never happens — so it suppresses total bases through OPPORTUNITY, not contact quality.',
build:(r)=>r.batter_k_pct*r.pitcher_k_pct,
},
];
/**
* Breaking-ball share of a pitcher's mix, from `pitch_mix` already ingested.
* Sliders/curves/sweepers/cutters vs fastballs — absent mix -> null, never 0.
// ── STEP 4 — COMBINED vs COUNTER (valid at this n; the gate is not) ─────
consth2h=rows.filter((r)=>r.skill!=null);
constys=h2h.map((r)=>r.won);
constbs=bootstrapDiff(h2h,'skill','champ');
console.log(JSON.stringify({
stat:STAT,
archetype_restriction:ARCH||'none (pooled)',
VALIDITY:contaminated
?'CONTAMINATED / DIRECTIONAL ONLY — statcast_aggregates now carries a single as-of date ('+freezeDate+') that is AFTER the settled games, so season profiles contain the games being predicted. These are NOT gate verdicts. statcast_history (new) makes point-in-time possible from tomorrow.'
:`CLEAN out-of-sample: profiles frozen ${freezeDate}; only game_date > ${freezeDate} scored`,
contaminated,
rows_scored:rows.length,
gate_spec:cv.VALIDATION_REQUIREMENTS,
bonferroni_tests:TESTS,
multiple_comparisons:{...mc,note:'denominator is DISTINCT hypotheses across the programme lifetime, not this session'},
n_gap_note:`the gate needs ${cv.VALIDATION_REQUIREMENTS.min_historical_instances} rows; this run has ${rows.length}`,
cluster_note:'76% of game pen-quality variance is WITHIN team, so the game is the honest cluster; team-clustered reported as the conservative sensitivity',
?`power ${spec?spec.power:0} < floor ${LODO_POWER_FLOOR} — this test would miss a real date-driven failure more than nine times in ten, so it can neither pass nor fail the stat`
:(exceedsCutoff
?`${reversals.length} reversals among ${informative.length} informative drops exceeds the cutoff of ${cutoff}`
:`${reversals.length} reversals among ${informative.length} informative drops is within the cutoff of ${cutoff}`),
console.log(` American cents — median ${r.cents.median} | p75 ${r.cents.p75} | p90 ${r.cents.p90} | max ${r.cents.max}`);
console.log(` implied-prob pt — median ${r.prob.median} | p75 ${r.prob.p75} | p90 ${r.prob.p90} | max ${r.prob.max}`);
console.log(` CROWN THRESHOLD (median ≥${CROWN_CENTS}c OR ≥${CROWN_PROB_PTS}pp): ${r.crownShips?'MET → crown MAY ship':'NOT met → crown does NOT ship'}`);
mechanism:'THE theorized carrier. Strikeouts need a pitcher who can miss bats AND a lineup that can be missed. An elite arm against a contact lineup and a modest arm against a whiff-prone one can produce the same count, so neither factor alone orders the props — the product should.',
mechanism:'ARCHETYPE-CONDITIONAL. Stuff should govern strikeouts more for a power arm than for a finesse arm, whose Ks come from chase and sequencing. Discipline 2 as a testable claim, with a categorical conditioner independent of whiff by construction.',
mechanism:'The finesse channel: expanding the zone only works against a lineup that will chase. Same shape as the stuff term, different mechanism, so it is tested separately rather than assumed to be the same effect.',
* THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the
* input is absent — an absent factor must leave the baseline untouched rather
* than nudge it toward some default.
*/
constFACTORS=[
{
key:'defense_by_direction',
needs:['spray_multiplier'],
entity:(r)=>`${r.player_key}|${r.opp}`,
mechanism:'CAUSALLY-CORRECT DEFENCE. Where the hitter puts the ball (pull/straight/oppo x ground/air) crossed with the OAA of the fielders actually standing in those zones, joined by handedness. Team-average failed the gate because it averages in five fielders who will never touch his ball.',
apply:(r)=>r.spray_multiplier,
},
{
key:'defense',
needs:['team_defense'],
entity:(r)=>r.opp,
mechanism:'A ball in play becomes a hit or an out partly by who is standing behind the pitcher. Should matter most where contact stays in the park.',
// More outs converted above average -> fewer hits.
mechanism:'Some parks turn outs into hits without producing runs — big outfields, high walls, deep gaps.',
apply:(r)=>r.park_factor,
caveat:'STAT_BASE maps hits -> run_base, so this is a RUN factor standing in for a HITS factor. A park that converts outs to hits without scoring is invisible to it.',
},
{
key:'platoon_severity',
needs:['platoon_severity_mult'],
entity:(r)=>r.player_key,
mechanism:"CAUSALLY-CORRECT PLATOON. The advantage is worth only what THIS hitter's measured split is worth, shrunk toward league by the smaller side's PA and refused outright below a floor. Flat handedness applies the same boost to a 63-point split and to none.",
apply:(r)=>r.platoon_severity_mult,
},
{
key:'platoon',
needs:['platoon_edge'],
entity:(r)=>r.player_key,
mechanism:'Handedness advantage — a hitter facing the opposite hand sees the ball better and hits it harder.',
apply:(r)=>(r.platoon_edge>0?1.06:0.96),
},
];
asyncfunctionmain(){
if(!SB_URL||!SB_KEY)thrownewError('SUPABASE_URL / service key required');
* THE FACTORS. Each returns a MULTIPLIER on the base rate, or null when the
* input is absent — an absent factor must leave the baseline untouched rather
* than nudge it toward some default.
*/
constFACTORS=[
{
key:'barrel_rate',
needs:['barrel_pct'],
entity:(r)=>r.player_key,
mechanism:'THE EXTRA-BASE SKILL ITSELF. A barrel is the exit-velocity and launch-angle combination that produces extra bases; it is the most direct expression of what total bases measures, where for hits it is largely irrelevant to whether a grounder finds a hole.',
// UNITS: fromStatcastRow returns barrel_pct as a FRACTION (0.06), not the
// 0-100 the raw table stores. Writing this against the percentage scale
// clamped every row to the maximum negative shift, which then "improved"
// Brier only by leaning on the counter's known global over-prediction.
mechanism:'Whether a struck ball becomes a double, clears the fence, or dies at the track. The atom reshapes HIT TYPE rather than P(hit), which is the only form that can express a total-bases effect.',
apply:(r)=>r.park_weather_ratio,
caveat:'venue-borne: replication caps at the number of distinct park readings, not the row count',
},
{
key:'platoon_severity',
needs:['platoon_severity_mult'],
entity:(r)=>r.player_key,
mechanism:"The hitter's OWN measured split, shrunk by the smaller side's plate appearances and refused below a floor.",
apply:(r)=>r.platoon_severity_mult,
},
];
asyncfunctionmain(){
if(!SB_URL||!SB_KEY)thrownewError('SUPABASE_URL / service key required');
premise_correction:'statModel.js and correlateValidator.js do not exist in this repo. The gate was implemented to the spec in src/services/python/blueprints/unconventional.py (VALIDATION_REQUIREMENTS); supplementSystems.test.js inlines its own validateFactor and imports no implementation.',
out_of_sample:`skill profiles frozen ${freezeDate}; only game_date > ${freezeDate} scored`,
mechanism:'Extra bases need BOTH conditions: hit hard AND hit in the air. A 105-mph ground ball is an out; a 25-degree popup is an out. Neither factor alone predicts bases, which is precisely why each may fail solo and the product may not.',
mechanism:'A hitter only realises his contact quality against a pitcher who permits contact quality. Elite suppression should attenuate a power bat; a contact-permitting arm should amplify it. The effect is conditional by construction.',
mechanism:'ARCHETYPE-CONDITIONAL. Barrels convert to extra bases for hitters whose lane is power; for a speed/contact profile the same barrel rate is a rarer event on a swing built for something else. This is Discipline 2 stated as a testable interaction. NOTE: it is currently UNTESTABLE — statcast rows carry no archetype label, and the barrel-relative proxy is an exact linear function of barrel_pct, so controlling for both components is rank-deficient. It needs a real archetype classification joined in.',
build:(r)=>r.batter_barrel_pct*r.archetype_power,
},
{
key:'batterK_x_pitcherK',
components:['batter_k_pct','pitcher_k_pct'],
mechanism:'Strikeout risk compounds multiplicatively (log5 is exactly this shape). A high-K bat against a high-K arm loses plate appearances to strikeouts, and a PA lost is a base opportunity that never happens — so it suppresses total bases through OPPORTUNITY, not contact quality.',
build:(r)=>r.batter_k_pct*r.pitcher_k_pct,
},
];
/** Latest settled game date in the pull — used to detect that the profile
* freeze now sits AFTER the data, i.e. no clean out-of-sample window exists. */
// ── STEP 4 — COMBINED vs COUNTER (valid at this n; the gate is not) ─────
consth2h=rows.filter((r)=>r.skill!=null);
constys=h2h.map((r)=>r.won);
constbs=bootstrapDiff(h2h,'skill','champ');
console.log(JSON.stringify({
stat:'total_bases',
VALIDITY:contaminated
?'CONTAMINATED / DIRECTIONAL ONLY — statcast_aggregates now carries a single as-of date ('+freezeDate+') that is AFTER the settled games, so season profiles contain the games being predicted. These are NOT gate verdicts. statcast_history (new) makes point-in-time possible from tomorrow.'
:`CLEAN out-of-sample: profiles frozen ${freezeDate}; only game_date > ${freezeDate} scored`,
contaminated,
rows_scored:rows.length,
gate_spec:cv.VALIDATION_REQUIREMENTS,
bonferroni_tests:TESTS,
multiple_comparisons:{...mc,note:'cumulative across the programme lifetime, not this session'},
n_gap_note:`the gate needs ${cv.VALIDATION_REQUIREMENTS.min_historical_instances} rows; this run has ${rows.length}`,
Inventory only. Every line cites a file or a query. Anything unverifiable is
marked UNKNOWN rather than asserted.
---
## PHASE 0 — Design / terminal
| item | state | evidence |
|---|---|---|
| **Scanner blue-boundary format** | **DONE** | `scan/page.tsx:715-717` — *"the blue boundary channel (`--priced-out`): line · BK odds · ◆ fair"* and *"renders NEUTRAL, not amber (blue-boundary law)"*. Amber survives only as an unrelated CTA (`:983`) and a status line (`:846`). |
| **MovementStrip** | **NOT-STARTED***(as named)* | No component by that name. Movement renders via `GradeShift.tsx`, `GameCard.tsx`, `MarketBreadth.tsx`, `FuturesBoard.tsx`. **UNKNOWN whether the spec wants a distinct strip or is satisfied by these.** |
| **THE WIRE** | **PARTIAL** | `components/vyndr/NewsWire.tsx` exists and is mounted — but only inside `ExploreHub.tsx`. No standalone wire surface. |
| **Book comparison** | **PARTIAL — built, unmounted** | `vyndr/BookComparisonPanel.tsx` + `components/BookComparison.tsx` both exist; grep of `web/src/app` returns **zero** mounting pages. Same built-but-unread class the reachability guard was written for; **not covered by that guard** (contract is grade fields only). |
| **Offseason hub** | **NOT-STARTED** | No component, no route. |
| **Terminal** | **DONE — deliberately retired** | `app/terminal/page.tsx` redirects to `/dashboard`; layouts preserved unrouted in `components/intel/TerminalTemplates.tsx` (S57). Not debt. |
| **Screen conversion (S36–39 arc)** | **DONE except one** | 47 route dirs under `app/`. `RouteStub` survives in exactly **one** file: `app/notifications/page.tsx`. Every other screen has a real page. |
**Sports surfaces that exist as pages:**`soccer` (414 lines), `desk` (160),
`intelligence` (160), `marketplace` (123). No `fight` route.
| **MLB · pitchers** | **SCAFFOLDED** | `model/pitcherEngine.js` (261 lines, own FLAME/SCALPEL/SINKER archetypes) — **read by no serving code** (grep: zero non-test consumers). Strikeouts n=57 settled vs a 500 gate. Base rate repaired by the shared `getStatRows` MLB fix. |
| **WNBA** | **SCAFFOLDED** | Archetypes exist; settles via ESPN box scores (376 rows historically); no factor model. Base-rate path fixed this session. |
| **NBA** | **SCAFFOLDED, dormant** | Archetypes exist; offline (Python service down, off-season). Base-rate paths fixed dormant at `55b210c`. |
| **Soccer** | **SCAFFOLDED** | 6 archetypes, a real `/soccer` page, feature extractor exists. No factor model, no settled outcomes. |
**Nothing but MLB batters is MODELED.** Everything else is scaffolding with a
sound base rate and no proven factors.
---
## PHASE 2 — Wiring / data-integrity debt
### `edge_pct` — the flag is half right, and the diagnosis was wrong
**Not a scale bug.**`analyzeViaEngine1.edgePctFor` computes
`(model − line) / line`, signed by direction — arithmetically correct. It
*explodes on small lines*: a 5.5 projection against a 0.5 line is a legitimate
1000%. Live top values: **900, 860, 700, 700, 660** on **68,364 rows**.
**Consuming surfaces: 15 backend files + 10 frontend files.** But
`edge_pct` itself reaches a user only through `alt_lines` typing in
`scan/page.tsx:64` — the grade card's `edge` is computed independently in
`gradeAdapter` via `computeEdge`. So the blast radius is smaller than the file
count implies.
**`ev_pct` DOES render** — 5 frontend files (`PriceTriplet`, `GradeResultCard`,
`LiveHeroProp`, scan page + route). The "ev_pct renders nowhere" flag is **stale**.
45,125 rows carry it.
**Classification: PARTIAL — a real display defect (a 900% edge is not a sentence
we can defend), not a broken computation.**
### `opp_rank_stat`
`refreshTeamStats` is wired into `snapshotService:361`. **UNKNOWN whether it
populates in production** — not verifiable from the repo, needs a live probe.
### The raised standard. Commit to /specs/DESIGN-SPEC.md. Supersedes v1. Governs the design train and every future surface. Built from: cross-vertical research (Bloomberg/Linear/Superhuman/Stripe), the live design audit (22 findings), and the founder's realism reference.
### Governed by MASTERMIND-BUILD-STANDARD (tokenized, tested, no one-off CSS). Every rule here is buildable — a token, a component, a motion primitive, or an asset.
## THE ONE LINE
The substance is already here — the record (55-19, 74% hit), the archetype system, real live data. v2's entire job: **stop the presentation from failing the substance.** VYNDR must feel like a $1M/month terminal that costs $44.99 — the gut reaction "how is this only $44.99" IS the conversion event.
---
## PART 0 — GOVERNING LAWS (from founder + research)
1.**ALIVE IS PUNCTUATED STILLNESS** — mostly still, meaningful pulses. Not constant motion. Fast motion reads cheap; fast RESPONSE reads expensive.
2.**PREMIUM = SPEED + RESTRAINT, NOT RICHNESS** — Bloomberg's moat is zero-latency; Linear feels expensive via sub-200ms + optimistic UI; Superhuman is "confident, not loud." Perceived instantaneity is the #1 premium lever.
3.**THE BLEND** — peak-ESPN swagger (weighted motion that ARRIVES then RESTS; one hero per screen) + Bloomberg density (dense but RANKED, effortless, instant) + VYNDR's existing glitch/signal-green signature (keepers — evolve, never replace).
4.**REALISM** — entities render as THEMSELVES: real team logos/colors, player identity, sportsbook wordmarks. Never a gray string.
5.**DENSITY IS UNLIMITED UNDER FIXED GRAMMAR** — the founder LIKES density; the failure is dense-and-flat. Rank it or it's noise.
6.**THE GAP IS THE PRODUCT** — design every surface so the price feels impossibly low for the experience.
---
## PART 1 — THE COLOR CONTRACT (fixes audit #3,#6,#9,#15; enforce as QA tests)
- **Signal-green #00D4A0 = ONE meaning: edge / active / A-tier / primary CTA.** Own it like Linear owns purple. It is VYNDR's signature color — unmistakable, never diluted.
- **Never green for:** decorative backgrounds (kill the green-flooded Intelligence panel #15), "live" status while stale (#6), two competing CTAs (#9), or archetype colors that are green-family (dedupe BRUSH/WORKHORSE off the signal hue).
- **Edge / CLV colored by SIGN:** negative = muted red, positive = signal-green. A -33.3% edge must NEVER render green (#3).
- **Grades colored by TIER:** A/A+ signal-green, B neutral-bright, C muted, D/F red.
- **GLOW = A/A+ ONLY.** A glowing C devalues the cue (#3). Glow is the scarcest signal in the system.
- **Muted default, full-saturation only on what matters** (the Linear pattern, audit-validated): most text/chrome at 40–60% opacity; saturation reserved for the one thing that matters per surface.
## PART 2 — THE ENTITY LAYER (fixes #10,#11,#12 — founder's #1 priority; one system, swapped everywhere)
A single asset/rendering system, cached, reused on every surface:
- **Teams:** real logo + primary/secondary colors. Game cards, grade cards, team hubs get the logo and a team-colored edge/accent. Never "MIL @ PIT" as bare text.
- **Players:** real headshot; fallback = team-colored monogram (not the gray silhouette). Always paired with position + status chip (CONFIRMED / OUT / GTD / NOT-IN-LINEUP) as a consistent identity block. Team hub already does position — propagate.
- **Sportsbooks:** real wordmarks in brand color (DraftKings, FanDuel, BetMGM, etc.), never lowercase text. Best-price gets a clear semantic highlight.
- **Reference standard:** the founder's admired game-strip — real logos, real O/U, park, context — executed to LEGIBLE premium (the reference's own cramped table is what VYNDR does BETTER: same density, real hierarchy).
## PART 3 — THE DATA-DISPLAY STANDARD (fixes #13,#14,#16,#19,#21 — the "horrid" core)
Every stat/line/player/game is a designed component, not a cell. Rules:
- **One bold hero figure among muted context.** The number that matters is large + mono + tabular; everything else demoted. The failure (#14,#19): every value identical weight.
- **Stats carry context (the lens):** never a bare number — value + rank + trend + matchup. ".312" is trash; ".312 · 4th · ▲ up from .287 · .340 vs this SP" is intelligence.
- **Collapse dead repetition (#14):** six rows of "Grades post ~6:00 PM ET" → one compact "6 props pending · grade 6 PM" line. High real estate must carry high information.
- **Rankings must vary (#13):** if a leaderboard's rows are identical (all doubles-U0.5-55%-B), the ranking is meaningless — surface the varying signal (edge, confidence, trend) and close dead-space gaps with context. Tabular mono, aligned.
- **Charts that are flat should SAY so (#16):** CLV 73/74 flat → "CLV flat — we grade the outcome, not the close" in one confident line, not a histogram that reads as broken. The best data (55-19, 74%) must look the most premium, never errored.
- **No naked "—" fields (#15,#21):** populate or hide. A Desk billing page with blank STATUS, an Intelligence panel with four dashes — both read unfinished.
- **Legends where needed:** the last-10 dot strip (●○○●) needs a hit/miss legend.
- **Kill loading walls (#5):** no "Loading the slate…" / "Assembling the pack…" text. Layout-matched skeletons only. FIX the React #418 hydration mismatch — server HTML must not be discarded and re-rendered.
- **Ticker = punctuated stillness (#7):** rests on each item long enough to READ (≥4s hold, or a static ranked "top moves" strip that pulses only on change). ONE animated element in the header zone — not ticker + heartbeat + counter stacked.
- **Motion categories:** IDLE (subtle proof-of-life — SYNC tick, one live-dot), TRANSITION (sub-200ms weighted), DATA-UPDATE (300ms green pulse on change — the terminal reacting), REVEAL (grade declassify, the Settle — staged, rare, deliberate), GLITCH (wordmark RGB-split on load, header scanline — chrome only, NEVER on data). All tokenized (--motion-*), prefers-reduced-motion honored globally.
- **The peak-ESPN rule:** motion ARRIVES with weight, then RESTS to be read. Swagger in the chrome, signal always pristine.
## PART 5 — ARCHETYPE SYSTEM (the Rosetta stone — propagate, don't reinvent)
The team hub already renders archetypes as an ownable visual language: glyph + color chip (★ALPHA, ▲BOMBER, »GHOST, etc.). This is the strongest asset on the site. v2:
- **Propagate the glyph+chip to EVERY surface** — dashboard, grade reveal, /desk, streaks, ledger, Data Brief (currently plain text / JSON in most). One component, everywhere.
- **Dedupe archetype colors off signal-green** (BRUSH, WORKHORSE currently green-family — shift them so #00D4A0 stays uniquely "edge").
- **Each archetype gets its one-line meaning on-page** where it leads (player page, grade reveal), expander for the long form.
## PART 6 — CONVERSION ARCHITECTURE (design that produces the million)
- **Never an empty hero (#1):** "Top grades tonight" must fall back to yesterday's A-tier settled receipts (you have 55-19) — first paint ALWAYS proves the model.
- **Desk $44.99 = the hero tier (#8):** rebuild pricing so the premium tier carries the "$1M terminal" story — real feature visuals, balanced layout (kill the dead right-half), a CTA at least as strong as Analyst's. Stop selling down.
- **Landing 5-second test:** one sharp positioning line about how it FEELS + live proof above the fold (real grades, the record). The product working IS the hero. Not features, not animation.
- **Build the missing /u profile (#4):** the #1 shareable surface doesn't exist — a discoverable, designed, CLV-verified public profile (Kev's partner-pitch ammo).
- **Fix scan→ledger persistence (#17):** a completed read must land in the ledger instantly — the "Every grade, no hiding" promise breaking is a conversion-killer.
- **Growth prompts don't occlude the product (#9):** the notify/PWA card shows once, dismissible, one primary CTA.
## PART 7 — SCREENSHOT-FIRST PRIORITY (the timeline is customer #1)
Design attention follows screenshot value: **P0 billboards** (Settle card, grade reveal, STREAKS row, /u profile, OG images) get pixel-level craft, tested at social crop ratios · **P1** landing/dashboard/pricing · **P2** slate rows/scan/explore/player/team · **P3** settings/onboarding. The STREAKS row (#2) is a billboard: real player image, streak-length as mono hero, muted "built vs" context, one severity accent — a Bloomberg alert, not a log line.
## PART 8 — EMPTY/ERROR SYSTEM (fixes #20)
The 404 page is already on-brand (glitch, "TRANSMISSION INTERRUPTED", correct CTA hierarchy). Make it the quality bar: ONE designed empty/error system reused everywhere. Kill the bare-red "Team not found"; unify the three offseason-state voices into one.
---
## THE DESIGN TRAIN (build sequence — sessions like the A1 board)
- **DS0 — Entity Layer** (Part 2): logos/colors/headshots/wordmarks system, swapped into every surface. Biggest single transformation; founder's #1.
- **DS2 — Dashboard Slate Rebuild** (the named #1 rebuild): real entities + one hero per card + collapse pending-filler + never-empty hero. Inherit the team-hub bar.
- **DS3 — Color Contract** (Part 1): green = one meaning, edge/CLV by sign, glow A-tier only, re-skin green-flooded panels + double-green CTAs, dedupe archetype greens.
- **DS5 — Pricing + Motion + States**: Desk-as-hero pricing, ticker to punctuated stillness, unify empty/error to the 404 bar, propagate archetype glyphs everywhere (Part 5).
Each DS session: tokenized components, QA-locked (the color contract becomes failing tests), prefers-reduced-motion respected, acceptance evidence, ship on green. Same discipline as A1.
— DESIGN SPEC v2 · the presentation catches up to the substance —
| **proj-v1.1** (distribution ladder) | **94.2%** | *(forms the projection, not a nudge)* | 437 settled — aligned gap **0.252 vs 0.352**, concentrated in `hits`+`total_bases` |
All three orthogonal (r ≈ 0 vs projection, `p_win`, line and each other). Each
promotes ONLY on its own axis-filtered holdout, and ONLY if **reliability AND
resolution** improve. **This is time, not code.**
Both promote only on their own axis-filtered holdout, and only if **reliability
AND resolution** improve.
## 📍 STATE AS OF 2026-08-01 (Order Zero, measured on prod with the real key)
| finding | number | what it means |
|---|---|---|
| **MLB slate invisible to us** | **64.8%** | our own allow-list, not the feed — now widened for DISPLAY |
| **books/prop, MLB** | 3.61 feed → 0.57 after filter | the filter cost, quantified |
| **books/prop, WNBA** | **4.21** feed → 1.20 | **WNBA is BETTER covered than MLB** |
| **consensus ruler** | **MARKET, not SHARP — ⏳ PENDING-RECOVERY, not permanent** | `matchbook`/`polymarket` = 0%, but **`pinnacle` ran until 07-30** (103,940 captures) and stopped. **Do not enshrine as permanent** until PropLine answers — see `BLOCKERS.md` |
| **ruler delta** (consensus − incumbent) | MLB mean +1.50 pts, median 0, **17% of props move ≥5 pts** | rulers genuinely differ; "better" is unproven |
| **MLB isotonic `p_win`** | **DECIDED** — reliability **0.0846**, resolution **0.190**, holdout **n=125** | **PROVISIONAL label RETRACTED 2026-08-01.** Calibration is **ruler-independent** (`estimateProbability` never sees a price; the fit is p_win-vs-outcome). Replicated on a fresh later window, both metrics improved |
| **edge vs the ruler** | corr(edge, outcome) **−0.010** (v1) → **−0.022** (v2), n=200 · corr(**p_win**, outcome) **+0.26** | **Subtracting the market DESTROYS the signal.** The consensus ruler does not rescue edge: *differs ≠ better* |
| **exchange-inclusive ruler** | **UNTESTABLE on existing data** | exchange quotes were never stored (discarded until 2026-08-01). Becomes testable only as v2-era captures accrue |
| **WNBA** | **MODEL NOT BUILT YET** — held out, *not* failed | **CORRECTED 2026-08-01.** The −0.12 was **NBA-template machinery run on WNBA data**. WNBA has never had its own archetypes/variables/conditions — the "sport stubbed in on another sport's template" CLAUDE.md forbids. That is an **unbuilt model's expected failure, not a verdict on the sport.** Its own build is QUEUED, after MLB |
| 🔴 **pinnacle feed** | **0 captures since 2026-07-31** (103,940 in the prior 10 days) | a live regression; **we had a sharp anchor and lost it.** Not caused by our changes |
| **soccer** | **settles** — ~15 competitions, 30d | "grades into a void" is a **$19/mo Pro-tier** problem, not a data problem |
| **CLV + results feeds** | `/odds/closing` + `/movement`**redacted**; `/results`**403 `required_tier: hobby`**; `/exports/resolved-props`**403 `required_tier: pro`** | **verified on our keys** — plain tier exclusion, not a key or plan fault. **$9/mo** buys CLV + steam + results; **$19/mo** adds the 90-day settlement export |
| **books SERVED** | 5 → **13**; props rendered **546 → 2,780** (5.1×); mean **4.22** books/prop | **AGGREGATOR widening is LIVE.** Of 2,234 newly-visible props, **31.2% carry a real non-DFS price**; **68.8% are DFS-only** — shown, tagged, never a market |
| ✅ **S59 join invariant** | **ARMED** — root cause was `searchPlayer` returning `team: null` (`currentTeam` has no `name`); fail-safe: drops only on a positive not-in-game | `tests/unit/slateJoinInvariant.test.js` |
| ✅ **arch-v1 condition axes** | **BOTH NOW FIRING** — environment 84.6%, matchup 82.9% (`batter_own_split`). Was: env + matchup on 0/634 prod rows — `team` was null on 416/416, so the venue join had no key. **Environment FIXED** (joins on the game; resolves 105/120). **Matchup still dead** — needs opposing SP + both hands, three separate absences | `specs/arch-v1-axis-audit.md` |
| **opportunity_drift** | built, orthogonal (r≈0), live as a challenger on 142 rows/slate | verdict n-blocked; `scripts/opportunity-axis-holdout.sql` |
| **opportunity layer** | **NOT BUILT** — `ab_per_game` is display-only; engine1 has no usage factor; MLB batting order unavailable in every wired source | Step 0 stopped before wiring. Recommended instead: an `opportunity_drift` axis on the EXISTING `challengerProjection` (arch-v1). See `specs/connect-opportunity-step0.md` |
| **MLB board size** | **7 → 365 graded props** (52×) in 114s | the 25-cap discarded 95.7% of the slate. Raised to 500 on measured cost. Refusals were **43.8% deliberate policy suppression**, not a data gap |
| **model input** | **byte-identical** — 546 gradeable props, `v1_first_book` | `MODEL_BOOKS` gate in `dedupeProps` + `indexOddsProps`. Lifts only on the re-run |
| **accrual clock** | **sequential, post-completion** | see §11. Pre-completion data does not count and is never pooled |
**The honest framing:** widening books is an **AGGREGATOR** win. **It does not fix
the model.** Do not let the free-side win read as model progress.
> **HOW TO USE:** this supersedes ad-hoc re-derivation. Before any order, read the
> phase you're in. After any order, tick the item and add one line. **Do not
> re-audit anything marked KNOWN** — that redundancy is what this document exists
> to kill.
---
## 0. VERIFICATION LEDGER (what was re-checked in this pass)
**NOTHING was re-verified. No query was run.** Everything required is already
captured in 22 artifacts produced this session plus the canonical board. Per the
order's own clause — *"If everything needed is already in the artifacts, say so and
skip verification"* — this is that case.
**Taken as KNOWN (source in brackets):**
- Model architecture, all layers [`model-architecture-recovery-map.md`]
**KNOWN and load-bearing: settlement WORKS** (scheduler → `settleAllOutcomes` +
`settleAllLedgers`, 937+ settled, growing daily). **What is unreachable is the
user-output TAIL:** `/api/grading/resolve` has no caller, and its fanout holds
webPush/Telegram/Discord but **no share-card step and no recap**.
**D1** wire a trigger (or move the fanout into the settle pass) · **D2** share-card
generation + `/notifications` consent + result posts + recap · **D3 CLV flag
decision** — `clvCaptureReliable()` is *one env var*, and the pre-registered rule
stands: flip only if close_moved is a clear majority AND coverage is representative.
**🔴 Never wire `/api/grading/resolve` as a second settlement path — it double-counts.**
---
## E. SPORT BOUNDARY
Adding a sport is a **~10-file core edit** with four silent-failure modes
(MARKET_MAP → zero props; three stat whitelists → silent 400s; missing projection →
universal refusal; no settled feed → grades forever). **Collapse to a registry** so
a sport is a module. **Blocks all of A2 after MLB.**
---
## F. CHROME AUDIT — 11 items, 4 need a Desk session. **Runs when surfaces are stable, not before.**
---
# THE PHASES — 7 phases, ~18 orders
| phase | orders | contents | blocks |
|---|---|---|---|
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only — every other sport is NOT-BUILT, held out) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
> confirmation — richer than what we hand-built). `/odds/closing` and
> `/movement` are **REDACTED** on our tier (full structure, **zero prices**) —
> my first pass wrongly called them "works" on a non-empty body. `/results` and
> `/exports/resolved-props` are **403**.
>
> **But the settlements exist to be bought:** `/markets/resolution-summary`
> shows **soccer graded across ~15 competitions in 30 days** (MLS 41k, Liga MX
> 15k, Brasileirão 12k, UCL/Europa/Conference…). **"Soccer grades into a void"
> is a $19/mo Pro-tier problem, not a data problem.** NBA is absent because it
> is July — seasonal, not a coverage gap, and not inferable either way.
>
> **WNBA is NOT thin at the feed** — 4.21 books/prop vs MLB's 3.61. It was
> allow-list-starved exactly as MLB was. This removes one candidate explanation
> for its −0.12 result. **And that result is not a WNBA verdict anyway** — it
> measured NBA-template machinery on WNBA data. WNBA is **NOT BUILT YET**, held
> out until it gets its own model.
>
> **No sharp anchor exists for props:** `pinnacle`, `matchbook` and `polymarket`
> all measured **0%** on both sports. The consensus ruler is therefore a MARKET
> consensus, not a SHARP one — stated as a permanent limitation, not a milestone.
## 10.2 What a real DATA AGGREGATOR has that we don't
| capability | ours | gap |
|---|---|---|
| **Book breadth** | 5 admitted of **18 sent**; 4.41→1.50 books/prop after our own filter | **not a feed gap — one allow-list (10.1).** Core props are already 10–12 books wide |
| **True consensus / no-vig line** | single-book de-vig | **unblocked now**: median across reference books (exchanges + pinnacle), n≥2 or labelled fallback |
| **Historical odds archive** | **STARTED** — `closing_captures` 844k rows, `lock_lines` (033) new, in-grade history capped at **24 points** | no full open→close series per prop. This is what makes CLV and backtesting real |
| **Market breadth** | 11 live markets | the category ships 50+ (alt lines, combos, innings, quarters) |
| **Alt-line ladders from books** | we *compute* a ladder; we don't *ingest* the books' | users shop rungs |
| **`specs/design-vs-build-gap-audit.md`** | **A 61-item design-vs-build audit already exists** (2026-07-31, audited at `bf7c0a3`), every claim grep-verified, with a wave-ordered build plan. |
**The inventory this order asks for was largely already done on 2026-07-31.** The
useful work is reconciling it, not redoing it.
## PHASE 1 — per media surface: spec, and build state
| surface | design spec? | built state | evidence |
|---|---|---|---|
| **Offseason hub** | **YES — a full surface file** | **NOT-STARTED** | `Vyndr Offseason.dc.html` (119KB) defines hub home (NFL desktop + 390), an **NBA Summer-League variant with an `OUTLOOK ONLY / NOT GRADED` honesty block**, a season-long board (open→NOW→VYNDR triplet), a season-read reveal with `WHAT WOULD CHANGE THIS READ`, a news/outlook feed with row anatomy, and a **quiet-wire empty state**. Gap audit: F9/F10/F11, *"design complete, no blocker but big."* |
| **Article media** | **YES — surface S3** | **NOT-STARTED** | Hero template + **4 hero graphic archetypes** (line path / distribution / mark-at-scale / matchup card — *all generated data visuals, never stock*), inline figures (stat-callout triptych, comparison bars, pull quote — **one max**), **caption law: every figure names its data**, article card, OG 1200×630 (master already rasterised in `exports/`). Gap audit F5. Zero components built. |
| **THE WIRE** | **YES** | **BUILT** (E19) | Spec'd in `Vyndr System.dc.html`: timestamped entry, ~6s hold, tag coloured by meaning. Built as `vyndr/Ticker` on real `/api/ticker` exhaust. **`NewsWire.tsx` is a different thing** — the offseason news/outlook feed, mounted in ExploreHub. |
| **Movement strip** | **YES — E1, a named primitive** | **ABSENT** | *"steps not curves, green only when the move favours the read, FLAT = hairline + `FLAT · [N]D`"*, row 86×20 silent / full-width annotated on reveal. **This corrects my 2026-08-07 board, which listed it UNKNOWN / possibly-satisfied-by-GradeShift.** It is a specified primitive that does not exist. |
| **Share-card masters ×5** | **YES — E16** | **ABSENT as product** | settle 1080×1350, story 1080×1920, square 1080, X 1200×675, record 1080×1350, article OG 1200×630. PNG masters in `exports/`; `ShareCard.tsx` has **0 importers**. Gap audit blocks it on the resolution tail having no share-card generation step. |
@@ -23,8 +23,8 @@ graded prop inline) reads left → right in this fixed order:
| 3 stat+line+side | `TB O1.5` | white, mono — the subject of the row |
| 4 market context | best-price dot · movement chip (STEAM ▲ / VALUE ▲) | what the MARKET is doing, before what the model says |
| 5 model output | revision strikethrough (original grade) → current grade badge | history then present: `A̶ B` reads "was A, now B" |
| 6 outcome | settled chip (✓ HIT / ✕ MISS / PUSH + actual) | once settled, actions (slot 7) disappear — the bet is over |
| 7 actions | parlay `+` · BOOK IT ⟶ | suppressed when dead or settled |
| 6 outcome | live TRACKING mark (`1/2 TB · ▲6th` + progress bar + state chip), then settled chip (✓ HIT / ✕ MISS / PUSH + actual) | S11 amendment: live progress is a PROTO-OUTCOME and lives in this slot; the two are mutually exclusive (a settled prop shows only the settled chip). Once settled, actions (slot 7) disappear — the bet is over |
| 7 actions | parlay `+` · BOOK IT ⟶ | suppressed when live, dead or settled (the pre-game market for a locked line closes at first pitch) |
| 8 provenance | `Graded 2h ago at -115` | the receipt, always last |
**Sub-line (stat + market context, one line under the row, in this order):**
@@ -49,6 +49,13 @@ Red is never used for mere movement-against or a lower price — those are amber
the net move is TOWARD the graded side, amber when net AGAINST, dim when flat —
never red (nothing has settled).
**S11 amendment — live TRACKING marks (slot 6 proto-outcomes):** green =
already-cleared over (HIT ✓, filled), on-pace over, or an under still holding;
amber = NEEDS N or an under whose line has been passed. An IN-PROGRESS prop is
NEVER red — red stays reserved for the settled miss and the NOT-IN-LINEUP dead
read. The read is locked pre-game and never re-grades; every live card carries
"TRACKING — read locked pre-game" exactly once (specs/LIVE-TRACKING.md).
## 4. MARK LAW — the micro-vocabulary
- **● / ○ dot strip** — last 10 games vs TONIGHT'S locked line, newest first.
Filled green = that game's stat cleared the line; hollow dim = it didn't.
# The accrual watch — the program is idle on modeling, and that is correct
## PHASE 0 — live-surface integrity
| check | result |
|---|---|
| Grade bands not derived from the retired champion | **PASS** — `gradeBands` is required by *no* serving code. Built across several orders, never wired. No stale band can reach a user because none reaches a user at all. |
| No withdrawn-map leak | **PASS** — `CALIBRATION_DEPLOYED` is `[]` and the calibrate loop iterates it, so `calibrate()` is never called. The only `p_win_calibrated` assignment sits inside that empty loop. Served `p_win` is repaired-champion raw. |
| Refusals on the repaired reference | **PASS** — `projectionFor` reads `l20_avg`, which `mlbGameLogFeatures` now builds from `fullLog`. |
| Factors still fire post-repair | **PASS** — the repair moved the base they adjust, so sign was re-verified across it: at base 0.35/0.50/0.65/0.80, defence lowers, pitcher-contact raises, platoon raises at every point. Firing check only; **not** a lift re-measurement. |
### One defect found, and it was mine
The grade card rendered **"Last 20 games average: X"** from `l20_avg` — a field
that, after the repair, holds a **full-season** average. The number moved and the
label did not, so the surface asserted a window that no longer existed. Fixed to
for(constfoffactors)idx+=f.delta;// flat sum of ±1.0 / ±0.5
idx=clampIndex(Math.round(idx));
```
`grade_thresholds.json` is **not an input mapper in the JS path.** Nothing compares a
probability to those cutoffs. `engine1.js:29-36` reads the table *backwards* — it takes
the letter the index already produced and looks up that band's **midpoint** to
manufacture a confidence number.
**So `confidence` is a cosmetic re-encoding of the letter.** It carries zero information
beyond the letter and by construction can never disagree with it. There is **no
data-sufficiency penalty in the live path** — the one CLAUDE.md describes lives in
`mlbGrader.js:50-69`, which is DEAD CODE. (`computeFeatures.js:21` still carries a stale
comment claiming the adapter downgrades confidence; it does not.)
### Six of thirteen factors are wired to features nothing populates
| Dead factor | Δ | Why it never fires |
|---|---|---|
| `weak/top_opponent_defense` | **±1.0** | needs `opp_rank_stat` ← `team_stats:{sport}:{abbr}` ← **`refreshTeamStats` has ZERO production callers** (verified: only its own export + tests) |
| `consistency_elite/boom_bust` | **±1.0** | MLB consistency logs come from the same dead `gameLogService` path as Finding 2 |
`teamId`/`season_type`/`game_count_in_7d` through `gameContext`, give
`safeGetConsistency` the same MLB branch as step 1. This restores ±3.0 of range and
makes A/D reachable **on merit**.
- **(b) Re-scale the index/thresholds** so the current narrow spread spans more
letters. **This is the tempting one and it is the wrong one** — it would mint A's
without adding a single bit of information, and every "A" would be a relabelled B.
It converts a visible limitation into an invisible lie.
**Recommend (a), explicitly reject (b).** If (a) proves infeasible, the honest fallback
is to keep the two-letter output and stop advertising a scale we don't produce — not to
stretch the scale.
**Copy consequence, either way:** "A-RATED" appears on public surfaces and `AccuracyBadge`
for a grade the engine has never emitted. Until (a) lands, that copy is unsupported.
Open question for Kev: NBA/WNBA have no free game-log source on this path with Python
down. `espnStatsAdapter.getPlayerGameLog` (Wave 0) already solves exactly this for
`featureCache` — reusing it here is the obvious candidate, and costs no quota.
---
# RESOLUTION — Session 63 (shipped)
Kev's ruling: **(a) fix on merit, never (b) rescale.** Rescaling would mint A's
without adding information — a relabelled B marketed as an A, corrupting an
append-only ledger permanently. That option is permanently rejected.
## What shipped
| Fix | File | Effect |
|---|---|---|
| Normalized per-game rows for ALL sports | `featureCache.getStatRows` | Revives `p_win` → `ev_pct`, `kelly`, `model_odds`, `value`, hero v2. Also feeds consistency. |
| Rows wired into the grade path | `computeFeatures.safeGetConsistency` | One fetch per prop, shared by 3 starving consumers |
| `refreshTeamStats` called in production | `snapshotService.runSnapshot` | `opp_rank_stat` populated → the ±1.0 opponent factor can fire (it had ZERO callers) |
| `game_count_in_7d` derived from real logs | `computeFeatures` gameContext | `heavy_workload_7d` (−0.5) can fire |
| **L20 symmetry** | `engine1.computeFactors` | NEW `l20_contradicts_*`−1.0. There was no negative L20 path at all — a structural reason D was unreachable |
**Removing all three adjustments changes resolution by nothing on every stat, and
on rbi/runs it IMPROVES it.** Exactly one ablation anywhere has a CI excluding
zero — rbi home/away, and its sign says removing it makes the model **better**.
**So ~100% of the champion's resolution is `base + recency`: how often this
player has cleared THIS number lately.** That is the whole model. Everything else
is decoration.
### A correction to how we read last session's scoreboard
Pooled across stats the champion resolves **0.46**; per stat it is **0.196
(hits)** to **0.499 (rbi)**. Pooling stats with different base rates *inflates*
correlation, because p_win varies across stats in the same direction as the true
base rate. **0.46 is a pooling artifact and should not be quoted as the
champion's resolution.** The paired *differences* in the scoreboard remain valid
(champion and challenger were pooled identically); only the absolute level was
inflated.
## 3. Do the challengers have it, or dilute it? (STEP 3)
| challenger | uses the base-frequency signal? | verdict |
|---|---|---|
| **proj-v1.1 ladder** | **No — it replaces it.** Fits a rate + NB distribution instead of counting frequency at THIS line | **DILUTING.** Measured reliably worse (−0.0301, CI excludes 0). It discards the one thing that works in favour of a lossier route to the same question |
| **hits-v1** | No — same substitution, binomial instead of NB | **DILUTING.** Refuted (−0.022, CI excludes 0) |
| **arch-v1 · environment** | Adds park/weather, which the champion ignores | **DILUTING.** n=871, −0.0028, CI includes 0 — movement without information |
| **arch-v1 · opportunity** | Uses `opportunity_drift` — **the one feature with repeated residual signal** | **HAS THE FEATURE, WRONG IMPLEMENTATION** (see §4) |
| **arch-v1 · matchup** | — | STILL PENDING (rows settle after ET midnight) |
| **contact-v1** | Statcast contact quality; not in the retained vector | No evidence either way (n=1,055, CI includes 0) |
## 4. The missing-feature test — one real lead, already in our hands
Correlation of each **unused** feature with the champion's residual (`won −
p_win`), per stat, bootstrap CI.
**Multiple-comparisons discipline first:** 14 features × 5 stats = 70 tests at
α=.05, so **3–4 CI-excludes-zero results are expected by chance.** Six appeared.
A single hit is noise. **Only a feature that repeats across independent stats is
Per the product doctrine, calibration is **half** the success criterion — "does
60% mean 60%?" Right now 58.5% means 55.0%, and on total_bases 53.4% means 45.8%.
**This is fixable with no new data at all.**
## 7. VERDICT PER STAT (STEP 4)
| stat | n | resolution | verdict |
|---|---|---|---|
| **hits** | 578 | 0.196 | **DILUTION + CALIBRATION.** Adjustments contribute nothing; one real lead (`opportunity_drift`) that we already compute and implement wrongly; +4.4pt over-prediction |
| pen **archetype** | did NOT prove (0.5669 vs 0.5309 modal baseline, corrected interval spanning zero) — **excluded**, building it in would chain on an unproven link |
| hitter **approach identity** ("fastball-hunter", "finesse-vulnerable") | **does not exist** in this registry. MLB batter archetypes are BOMBER / GHOST / TORCH / BRUSH / DRIVER / FLEX / ALPHA / HYBRID / CATALYST. Inventing an identity to condition on is the fabrication the gate exists to catch |
A hitter power/contact split derived from the sequence data itself was tested as a
**separate gated addition** rather than assumed into the main effect. Neither half
proved (power −0.0002, contact −0.0001, both intervals spanning zero).
---
## The gate — two-part, on the concentrated subset, 114 cumulative tests
| subset | n | games | mean shift | Brier Δ | CI | verdict |
### Wave 6 of the wiring/data train. Net-new sport (MMA/UFC). Governs the combat build. Honest free v1; the full matchup-GRADE engine is a DEFERRED sub-wave. Build toward `specs/design-reference/vyndr-system.html` (FIGHTER A / VERDICT / FIGHTER B, GRAPPLER%/STRIKER% blend, tale-of-the-tape, MONEYLINE/DECISION/SUBMISSION/KO grades). Keep the live wordmark.
## DATA SEMANTICS (unchanged, restated for combat)
VYNDR never generates odds — ML/round-total values are REAL book numbers at a timestamp. Fighter records/physicals/style stats are REAL sourced facts. Model output (style edge, any grade) is always labeled MODEL. `Number(null)===0` is the trap — absent stat ⇒ absent, never 0. **Never ingest a fighter's photo/likeness** (same rule as headshots) — initials monogram only. Scraping fragility ⇒ degrade to absent, never fabricate.
## SCOPE — v1 (this wave) vs DEFERRED
**v1 ships:** fight-card discovery + tale-of-the-tape + style-blend archetypes + ML & round-total odds + a **style-edge VERDICT (a MODEL style read, explicitly NOT a settled grade)**.
**DEFERRED (own sub-wave, do NOT build now):** the matchup-GRADE engine that produces settled ML/method/round grades; method-of-victory & round & fighter props (data-limited on the free feed); combat outcome SETTLEMENT (no free box-score settle path yet — grades would stay `pending`); ufcstats.com scraping (needs a parser dep — CONFIRM before adding).
## DATA SOURCES (zero-out-of-pocket)
- **ESPN MMA (FREE, JSON, no auth — same family VYNDR already uses):**
- Fight cards / schedule: `site.api.espn.com/apis/site/v2/sports/mma/ufc/scoreboard` (date-pinned like the other sports).
- Event / fight detail + results + tale-of-the-tape: `.../mma/ufc/summary?event={id}` and the athlete endpoints (`athlete.id`, record, stance, reach, weight class). ESPN athlete id also gives a headshot via the entity layer's `a.espncdn.com/i/headshots/mma/...` pattern — but per the likeness rule, v1 uses initials monograms; wire the id but default to monogram.
- Depth caveat (honest): ESPN MMA striking/grappling granularity is THINNER than ufcstats. Style-blend is best-effort from what ESPN exposes (finish history, method-of-victory counts, takedown/strike splits where present); when a stat is absent, the blend says less — never invents.
- **The Odds API `mma_mixed_martial_arts` (ALREADY PAID — `ODDS_API_KEY`):** wire `oddsService.SPORT_KEYS.mma = 'mma_mixed_martial_arts'` + `MMA_MARKETS = ['h2h','totals']` (moneyline + round totals). Reuse `oddsNormalizer`. **PropLine has NO combat** → odds-api-only; combat does NOT enter the abundant-props path.
## COMBAT ARCHETYPE REGISTRY (PINNED — both `src/services/archetypeService.js` AND `web/src/lib/archetypes.js` use these EXACT names/colors/glyphs; a test asserts they match, same as other sports). Colors unique WITHIN combat; may reuse hues used in other sports (cross-sport reuse is fine).
Six styles. `classify('mma', fighterStats)` returns `{ primary, secondary|null, blend:[{archetype,weight}] }` (same contract as other sports) scored from finish-rate / method splits / takedown & strike tendencies.
| Name | Glyph | Color | One-liner (shown where it leads) |
|---|---|---|---|
| STRIKER | ✦ | `#E8703A` | Wins on the feet — volume + power at range. |
| GRAPPLER | ⊗ | `#2FA4E7` | Fight hits the mat on his terms — control + subs. |
| GRINDER | ▦ | `#B0883B` | Goes the distance, wins the rounds. |
STRIKER↔GRAPPLER is the primary range axis the mockup renders as the two blend bars; PRESSURE/COUNTER is tempo; FINISHER/GRINDER is the outcome tendency. A fighter is a BLEND (e.g. GRAPPLER 80% / STRIKER 30%). Discipline pedigree tags (Combat Sambo, Dagestan Wrestling, BJJ, Wrestling Base, Kickboxing, Muay Thai, Boxing) are **verifiable credentials**, rendered separately from the archetype blend — VERIFIABLE ONLY, mark unknowns absent, never guess a fighter's base.
## STYLE-MATCHUP VERDICT (v1 — a MODEL read, not a grade)
`styleMatchup(fighterA, fighterB)` → a style-edge verdict (e.g. "GRAPPLER EDGE · Islam") from comparing the two blends + finish/defense tendencies. This is a **descriptive MODEL read**, labeled as such — NOT a settled grade, NOT an edge %, NO fabricated confidence. The mockup's CENTER VERDICT. Keep it honest: if the data is too thin to call, say "STYLES EVEN / INSUFFICIENT READ".
## CONFIG WIRING (the recurring 4-layer-desync trap — do all in sync)
-`src/config/statFilters.js` + `web/src/config/statFilters.ts` — combat stat categories (if any props surface).
- Combat is NOT added to `snapshotService.ACTIVE_SPORTS` in v1 (no settle path) — fight cards + odds + style card render on-read; grades stay OUT of the locked-snapshot loop until the deferred engine.
## FRONTEND (toward the mockup)
- **Tale-of-the-tape / head-to-head style card** — `web/src/components/vyndr/FightCard.tsx`: FIGHTER A (initials monogram, name, record e.g. 26–1, stance, reach) · GRAPPLER%/STRIKER% blend bars · discipline pedigree tags · archetype chip(s) via the shared `ArchetypeBadge` · CENTER VERDICT (style edge) · FIGHTER B mirror · a grades/odds row (MONEYLINE + round total from the odds feed; method/KO/SUB shown as "— data-limited" honest placeholders, NOT fabricated). Two fighters side-by-side (NOT the player-strip row grammar). Mono data, no glitch on data.
- **MMA badge/color** — the `#D4AF37` combat token already exists in `shareCards/tokens.js`; add a frontend MMA sport token + `SportBadge` support.
- **Route** — `web/src/app/fight/[id]/page.tsx` (server wrapper + client card) and/or a combat surface on the schedule; add Next proxies for the ESPN-MMA read (S25 rule). Self-hide / honest empty (reuse `EmptyState`) off-season.
- Sportsbook links via the existing `bookLinks.js`/affiliate layer unchanged.
## TESTS
-`archetypeService`/`archetypes.js` combat colors MATCH (extend the existing cross-file color test).
-`classify('mma', …)` blends correctly from fixtures; thin data → fewer/absent style claims (never fabricated).
-`styleMatchup` → honest "insufficient read" on thin data; a clear edge on divergent styles.
- combat adapter: defensive parse (ESPN shape → normalized fight card; unrecognized → empty, never throw); injectable, NO network.
- FightCard self-hides / honest-empty; renders tale-of-the-tape mono; method/round shown as data-limited, not fabricated.
## OUT OF SCOPE REMINDER
No settled grades, no method/round/props board, no fighter photos, no scraping dep — all DEFERRED and flagged in-UI as data-limited rather than promised. The cheapest honest v1 is the goal: styles + tape + ML/round-totals, beautifully rendered.
`Vyndr Scanner States.dc.html` — spec for the two scanner states Code flagged missing (converges at the scanner-nudge build order).
- **S6 scan empty (marketless scan)**: read-card frame, ZERO market numbers (no empty triplet frame — it doesn't render), chip `NO MARKET` in the blue boundary channel (`rgba(106,147,200,.12)` bg / `.34` border / text `#8FB2DE` — same tokens as PRICED OUT), dashed blue void box (refusal frame shape, blue not red: nothing failed), copy "no live market for this line" (never "coming soon"). Path forward built in: green primary CTA to tonight's board with live priced-prop count; player-specific path (`HERBERT · 3 PRICED PROPS ▸`) leads when the player is on the slate. Mobile 390: both paths 44px.
- **S7 priced-line display (the nudge)**: strip under the scan field showing ONLY genuinely-priced lines per exact player+stat. Case A `NONE PRICED` is the DEFAULT/common case — quiet dashed blue box + "we price these for [player]" stat chips (with amber fair previews). Case B priced: 44px rows `o 27.5 · BOOK −114 · ◆ FAIR −105 · OPEN READ ▸`, one tap to the real triplet. Case C off-slate: text-tokens-only fact line, no blue, no chips, stops. Typed-line mismatch: query text preserved, `o 30.5 ISN'T PRICED` blue fact line (not an error), nearest priced line = same player+stat only.
- New color law: **blue #8FB2DE = the boundary channel** — the system being honest about what it can't hand you (edge priced out, no market, line not priced). Distinct from red REFUSAL (model decision) and amber QUARANTINE (model leg suppressed).
Two new design files, same package format. System law unchanged (tokens, grade colors, glow-A-only, THE VERB IS "READ").
-`Vyndr Offseason.dc.html` — Surface 1. Hub home (NFL desktop + 390, NBA Summer-League variant with OUTLOOK ONLY / NOT GRADED honesty block), season-long board (open→NOW→VYNDR triplet — open `#707080`, now `#F0F0F0` 700, model `#00D4A0` 800; ladder grouping: player header row + nested market lines at `padding-left:52px`), season read reveal (news-annotated movement, NEWS-DRIVEN FACTORS, WHAT WOULD CHANGE THIS READ kill conditions), news/outlook feed + row anatomy + quiet-wire empty state. Sport state lives IN the sport tab (`NFL · CAMP −5D`), never a separate offseason tab. Kickoff countdown is computed live from the real date (`Sep 10 2026`) — ambient, top-right, never the hero. New mobile tab bar: SLATE · EXPLORE · READ-FAB (50px green circle, void-colored V, `translateY(-14px)`, 6px void ring) · LEDGER · MORE.
- S2 line shopping: MOVEMENT STRIP primitive defined once (row 86×20 silent / reveal full-width annotated / laws: steps not curves, green only when move favors the read, FLAT = hairline + `FLAT · [N]D`). Book comparison: 7 BookChips (24px tile, brand-color mono wordmark, `image-slot` overlay for licensed logos — DK `#53D337`, FD `#1493FF`, MGM `#C4A45E`, CZR `#AD9660`, B365 `#FFE100`, ESPN BET `#D50A0A`, Fanatics `#2E6BFF`), best number crowned (green inset bar + BEST NUMBER chip + row tint `rgba(0,212,160,.04)`), disagreement mark (dots on an axis around the VYNDR reference line; SPLIT chip amber when spread wide). Push-to-book: primary green deep-link + secondary book buttons, geo-gated (buttons never render disabled — they don't render), affiliate microcopy verbatim. Geo-empty honest state included.
- S3 article media: hero template + 4 hero graphic archetypes (A line path · B distribution · C mark-at-scale · D matchup card — all generated data visuals, never stock), inline figures (stat callout triptych, comparison bars, pull quote — one max, caption law: every figure names its data), article card (thumb = hero graphic re-cropped), article OG 1200×630.
- S4 CLV: per-read chip `BEAT CLOSE +2.0` / `MISSED CLOSE −1.0` — color follows CLV sign, never the result; aggregate joins record line (`BEAT CLOSE 61%` + `AVG CLV`); no-data state names the tracking start date + N≥100 threshold, never a fake 0%. Calibration curve: headline claim ("When we say 70%, it hits 68.2%"), dots vs dashed perfect line, dot size = sample, buckets under N 30 render hollow/unlabeled.
- S5 The Report: hybrid email — dark billboard header + dark record band survive every client; light paper body (`#F7F6F2` / ink `#1A1A22` / green shifts to `#00A57D` on paper). 600px single column, system-font fallbacks, no images required to read. Signup module + /report archive (every issue shows its own day record).
- M2 crops in-file at 50%: square 1080×1080, story 1080×1920 (recomposed — hero letter scales, dot strip replaces table; dashed safe zones are design-time only, strip on export), X 1200×675, NEW record billboard 1080×1350 (30D record + tier hit-rate bars, C-tier honestly under .500). Exact-size PNG masters in `exports/` (square 1080×1080, story 1080×1920, X 1200×675, record 1080×1350, article OG 1200×630). Every S2–S5 surface also ships at 390 (Act 06 in the Intelligence file): market sheet (44px book rows, compact disagreement axis, sticky deep-link CTA), article view, ledger with per-read CLV chips + compact calibration card, /report signup + archive.
- Sample data is anchored to Jul 2026 reality; every number is a spec for what the live feed renders. Truth laws: no "tonight" off-slate, visible timestamps, honest empty states throughout.
Two design components, both open directly in a browser. All styling is inline (no stylesheets to port); every value below is already in the markup.
## Files
-`Vyndr Landing.dc.html` — desktop marketing landing: hero + live board proof, TIER RECORD calibration band, feature triptych, FREE / THE DESK pricing (no trial), wire footer.
-`assets/glyphs/` — all 83 archetype marks (74 display + 9 classifier-legacy) as individual SVGs (`currentColor`, 24-grid) + `MANIFEST.md` (name → hex → file). Generated from `glyphDefs()`; wire these into `archetypes.js`.
-`exports/` — rasterized share cards: `settle-1080x1350.png` (exact) and `grade-reveal-og.png` (1080×566, same 1.9:1 ratio as 1200×630 OG).
-`image-slot.js` + `.image-slots.state.json` — drag-and-drop real-asset targets (player headshot, streaks rows, combat fighters) layered over monogram fallbacks; empty chrome hidden via `image-slot::part(empty){opacity:0}`. Design-time tooling, not product code.
- Logo mark: 32-grid rounded tile, "signal V" — left stroke white, right stroke `#00D4A0` with dot terminal at (24,9). SVG inline in both files.
## Glyph library
`glyphDefs()` in the desktop file's logic class: 74 marks keyed by archetype name — `{c: hex, g: svg-inner-html}`, 24-grid, duotone (fill-opacity .16–.2 base + 2–2.6 stroke). Colors deduped off signal-green/amber/red. Lift directly into `archetypes.js`. Blend display: primary chip full color + supporting chip muted (opacity .7–.85).
## Behavior (all implemented in logic classes)
- Reaction primitive: `nudge(cell, up)` — value updates, `rx-up`/`rx-down` background flash (.75s ease-out), edge recolors by sign. Fire on data events only, never idle loops.
- Boot: rows stagger in (`vy-rowin`, 60ms steps), hero arrives last (`vy-arrive`), reactions gated ~1.5s.
- THE WIRE: one timestamped entry, holds ~6s (`vy-wirein`), tag colored by meaning.
- ⌘K palette, row-hover rationale, IntersectionObserver reveal + rail highlight — all in `Vyndr System.dc.html`.
## Laws
Alive = punctuated stillness (react to truth, then rest). Dense but RANKED — one bold mono hero figure among muted context; dim ranks progressively (1 / .86 / .64 / .48). One animated hero per zone. Real entities always (team-colored monograms until licensed assets drop in — headshot slots marked). No trial: free tier is the trial. Premium is produced, never claimed.
## Rev 3 (Jul 16)
- Matchup lines everywhere now carry inline team-gradient chips (10-12px, before each abbr) — recognition without weight; chips sit inside rows so the ranked opacity ramp dims them too. Pattern: `<span style="display:inline-block;width:10px;height:10px;border-radius:3px;background:linear-gradient(135deg,<c1>,<c2>);vertical-align:-1.5px;margin-right:3px"></span>ABR`. Swap for licensed logo imgs at the same size when assets land.
- 9 classifier-legacy marks added to `glyphDefs()` + `assets/glyphs/`: BRUSH #E0B84A, WHIFF #E86A6A, CONNECTOR #9AB0C4, DISTRIBUTOR #7AB8D8, FASTBREAK #4AA0E8, FLEX #A08AC8, HYBRID #C88AB0, SWITCH #C0B08A, SWITCHBOARD #90A0E8. Same 24-grid duotone language; colors deduped off signal-green/amber/red.
File diff suppressed because it is too large
Load Diff
Some files were not shown because too many files have changed in this diff
Show More
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.