The phases counted unbuilt code. This section names what is missing for VYNDR to do what it claims, including the parts that are not builds. THE CENTRAL GAP: there is no demonstrated edge yet. Every measurement this session returned null, negative or unproven — served grade r~0.005 and inverted; all three p_win-vs-fair_prob formulations negative on both sports and both splits; p_win alone on MLB holdout p~0.07; WNBA negative; CLV null by guard; ROI-by-grade likely an artifact. The product's core claim is not currently supported by our own data, and building all 23 orders without closing this leaves a well-built product that does not do the thing it sells. What closes it is sample and honest iteration, not code — roughly 6-10 weeks at the current accrual, a clock engineering cannot shorten and that must not be faked. Also named: the projection is thin (l5/l20 + opponent rank + rest + usage, with similarity/archetypes/conditions/Bayesian all built and disconnected, so connecting them is a hypothesis not a guarantee); it is a one-sport product today (NBA and soccer do not even settle); there are 3 users and 0 paid so nothing is validated by usage; there is NO distribution path at all, which appears in no phase and belongs on the board as its own track; the last mile is unclosed (push-to-book is a teaser, no affiliate live); and operational fragility remains (single box, two-sport settlement, three credentials flagged including a Stripe live key that transited a transcript, no staging). The honest summary: the truth infrastructure is genuinely well built and this codebase does not lie about what it knows. What is not yet true is that the model beats the market — not disproven, unmeasured at adequate n. The finish line is 23 orders PLUS a verdict from accrued data we cannot rush, and the discipline to report that verdict honestly if it says the edge is not there. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
13 KiB
VYNDR — MASTER PLAN
Single source of truth. Sessions EXECUTE against this and UPDATE it in place. Created 2026-07-31 by consolidation. Nothing built in this pass.
HOW TO USE: this supersedes ad-hoc re-derivation. Before any order, read the phase you're in. After any order, tick the item and add one line. Do not re-audit anything marked KNOWN — that redundancy is what this document exists to kill.
0. VERIFICATION LEDGER (what was re-checked in this pass)
NOTHING was re-verified. No query was run. Everything required is already captured in 22 artifacts produced this session plus the canonical board. Per the order's own clause — "If everything needed is already in the artifacts, say so and skip verification" — this is that case.
Taken as KNOWN (source in brackets):
- Model architecture, all layers [
model-architecture-recovery-map.md] - Grade↔outcome correlations, collapse cost [
full-output-grade-mapping.md] - p_win calibration + holdout verdicts [
grade-diagnostic-t0.md,pwin-recalibration-holdout.md] - Market-relative edge inversion [
grade-fix-part1-investigation.md] - Design implemented-vs-designed, 61 items [
design-vs-build-gap-audit.md] - Surface states, orphans, waves [
pre-audit-status-pull.md,incomplete-surface-triage.md] - Resolution pipeline true state [
resolution-and-clv-investigation.md,wave3-*.md] - Tier/monetization + founder mechanism [
tier-structure-pull.md,tier-redesign-spec.md,build2-review-zero-report.md] - Sport-boundary cost [
model-architecture-recovery-map.md§Phase 2]
Genuinely OPEN → carried as explicit unknowns (not verified because they need a build or a decision, not a query): sport order (Kev's call, §A2); board-reasoning gating (a/b/c, §C2); CLV flag decision (§D3); the ~70 undefined team colours (§B).
A. PER-SPORT MODELS — Phillips doctrine: each sport is its OWN model
A1. MLB layer stack (the reference module), bottom-up
| # | layer | state | note |
|---|---|---|---|
| 1 | Data/feeds | BUILT | statsapi free+unlimited; statcast; park; weather; probables. Settles end-to-end. |
| 2 | Similarity (comparable instances) | BUILT · NOT WIRED · NOT DEPLOYED | python/utils/similarity.py; grade path skips to season/recent averages |
| 3 | Archetypes (batter + pitcher) | BUILT · display only | classifies + renders; does NOT feed the grade |
| 4 | Variable weights | PARTIAL/ASSUMED | engine1's flat ±1.0/±0.5 deltas are hand-set, never fitted |
| 5 | Conditions (park/weather/platoon/arsenal) | BUILT · CHALLENGER ONLY | rides env_*/challenger_*; "measured, never served" |
| 6 | Bayesian inference | BUILT · NOT WIRED · NOT DEPLOYED | python/utils/bayesian.py, 320 ln, genuinely sport-agnostic math |
| 7 | Calibration | MEASURED, NOT APPLIED | isotonic qualifies on holdout (rel .1038→.0939, res .139→.123, n=119) |
| 8 | Grade ladder | BROKEN | letter is a factor index, r≈0.005, inverted (B 52.4% < C 56.9%) |
The through-line: layers 2, 3, 5, 6 are built and not connected; layer 8 is connected and meaningless. MLB's fix is connection, not construction.
A2. Sport order (OPEN — Kev decides)
Proposed by readiness × clock: 1) MLB (reference, only qualifying model) → 2) CFB (has a <30-day clock; soft-market thesis) → 3) NFL → 4) NBA → 5) CBB → 6) WNBA re-attempt (abstains today) → 7) soccer (quota-blocked). Each gets the same 8-layer template. No sport is abandoned — abstention is a state, not a verdict.
B. DESIGN IMPLEMENTATION — 61 items catalogued
BUILT-TO-SPEC 20 · DRIFTED 7 · PARTIAL 16 · ABSENT 18. Ordered wire-in: D1-A done (6 combat glyphs, boundary-blue completed, reaction primitives, READ-FAB). D1-finish done (rationale/reveal/chips modules — built, NOT mounted). Remaining: D1-close (mount — blocked on a ROW-GRAMMAR slot amendment + §C2) · D1-B (45 unwired glyphs + the 41-vs-74 archetype scope call) · S2 primitive set (movement strip, crown, disagreement axis, SPLIT) · S3 article media · The Report email + archive · Offseason artboards · team colours: only 10 of ~80 defined — the rest render honest-neutral until a real source exists.
C. SURFACES
C1 — done: Wave 1 wiring · /compare · /record · Build-1 gate.
C2 — OPEN DECISION (blocks D1-close): board reasoning is served ungated while
tiers.js declares reasoning_visible:false. Options (a) gate it, (b) accept as
free funnel, (c) leave unrendered.
C3 — remaining: /record has no nav link (the surface that justifies the
price) · /notifications · Offseason hub · /system · S3 media · /soccer
(quota) · share cards (blocked by D).
D. RESOLUTION PIPELINE → USER OUTPUT
KNOWN and load-bearing: settlement WORKS (scheduler → settleAllOutcomes +
settleAllLedgers, 937+ settled, growing daily). What is unreachable is the
user-output TAIL: /api/grading/resolve has no caller, and its fanout holds
webPush/Telegram/Discord but no share-card step and no recap.
D1 wire a trigger (or move the fanout into the settle pass) · D2 share-card
generation + /notifications consent + result posts + recap · D3 CLV flag
decision — clvCaptureReliable() is one env var, and the pre-registered rule
stands: flip only if close_moved is a clear majority AND coverage is representative.
🔴 Never wire /api/grading/resolve as a second settlement path — it double-counts.
E. SPORT BOUNDARY
Adding a sport is a ~10-file core edit with four silent-failure modes (MARKET_MAP → zero props; three stat whitelists → silent 400s; missing projection → universal refusal; no settled feed → grades forever). Collapse to a registry so a sport is a module. Blocks all of A2 after MLB.
F. CHROME AUDIT — 11 items, 4 need a Desk session. Runs when surfaces are stable, not before.
THE PHASES — 7 phases, ~18 orders
| phase | orders | contents | blocks |
|---|---|---|---|
| 1. MLB model truth | 4 | promote isotonic p_win (MLB only, WNBA abstains) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
| 2. Resolution tail | 3 | trigger · share cards + notifications + posts + recap · CLV flag decision | share cards, social proof |
| 3. Surfaces + design lane (parallel with 1-2) | 4 | C2 decision → D1-close mount · /record nav + remaining surfaces · D1-B glyphs/archetypes · S2 primitives |
Chrome audit |
| 4. Sport boundary | 2 | registry collapse · MLB re-expressed as the first module | all further sports |
| 5. Sport rollout | 1 per sport | CFB → NFL → NBA → CBB → WNBA retry → soccer, each on the 8-layer template | — |
| 6. Monetization finish | 2 | Stripe Phase-B live proof on first real signup · founder launch to the 3 existing users | — |
| 7. Chrome audit + hardening | 2 | the 11-item visual sweep · credential rotation + migration-drift reconciliation | ship |
Phases 1–4 and 6–7 = ~17 orders. Phase 5 = 1 order per sport (6 listed). Total ≈ 23 orders to the end state, of which ~11 are unblocked today.
DEFINITION OF DONE
VYNDR is complete when:
- MLB layers 1–8 are BUILT AND CONNECTED — similarity, archetypes, fitted weights, conditions and Bayesian all feed the grade; calibration applied; the ladder monotone (A>B>C, no inversion) and proven on held-out data.
- Every listed sport is finished on the same 8-layer template, or explicitly abstaining with its reason recorded — never silently absent.
- Design fully implemented — all 61 catalogued items BUILT-TO-SPEC.
- Every surface built, reachable and honest — no orphans, no live-but-not-honest surface, no dead component.
- Resolution pipeline live end-to-end — settle → share card / notification / post / recap, firing on a real settlement.
- Sport boundary is a registry — a new sport is a module, not a core edit.
- Chrome audit passed, logged-out and entitled.
- The record is publishable on its own terms — CLV either trustworthy-and- representative or honestly absent; no claim outruns its evidence.
The remaining work is finite and countable: ~23 orders across 7 phases.
STANDING LAWS (carried into every order)
Truth Law — absent beats wrong, no fabrication up or down · per-sport models, never a global engine · lookahead guard (lock-time fields only) · overfitting guard (fit one split, prove another) · never mint A's without new information · aggregate proof is free, itemized judgment is paid · one canonical founder flag · atomicity by unique index, never a count · cache-bust every post-deploy check · verify-after-write.
9. WHAT'S ACTUALLY MISSING FOR THIS TO WORK AS A PRODUCT
The phases above say what is UNBUILT. This says what is missing for VYNDR to genuinely do what it claims. Some of it is not a build, and one of it is not fixable by us at all. Written plainly because a plan that only counts code is the comfortable version.
9.1 🔴 THE CENTRAL ONE: there is no demonstrated edge yet
Every edge measurement this session came back null, negative, or unproven:
| measurement | result |
|---|---|
| served grade → outcome | r ≈ 0.005, and inverted (B 52.4% < C 56.9%) |
| p_win − fair_prob (3 formulations) | negative in all three, both sports, both splits |
| p_win alone, MLB, holdout | +0.165, p ≈ 0.07 — not significant |
| p_win alone, WNBA | negative — abstains |
| CLV / beat-close | null by guard — instrument not trustworthy |
| ROI by grade | likely an artifact of a meaningless letter |
The product's core claim — "our read is better than the market" — is not currently supported by our own data. Everything else in this plan is scaffolding around that. Building all 23 orders and not closing this leaves a beautifully-built product that doesn't do the one thing it sells.
What closes it: not code. Sample and honest iteration. The instrument fields are ~10 days old (442 rows). At ~90 decided MLB rows/week, a defensible verdict is 6–10 weeks out. That clock cannot be shortened by engineering, and any attempt to shorten it is the fabrication this whole session has been removing.
9.2 The projection — the actual engine — is thin and unvalidated
The grade's only real inputs today are l5/l20 averages, an opponent rank, rest and usage. Similarity, archetypes, park/weather/platoon and the Bayesian layer are all built and not connected. So VYNDR is currently a recent-form average wearing an intelligence system's clothes. Phase 1 connects them — but connecting them is a hypothesis, not a guarantee. They must each prove out on held-out data or be left disconnected honestly.
9.3 A one-sport product marketed as multi-sport
MLB is the only qualifying model. WNBA abstains on its own data. NBA and soccer don't even settle — they grade into a void. Until Phase 5, the honest framing is "an MLB product with other sports in development." The site should not imply otherwise.
9.4 No customers, therefore no feedback loop
3 users, 0 paid. The founder mechanism is built and race-proven, /record
exists, the gate works — and none of it has met a real user. Nothing here is
validated by usage: not the price, not the tier line, not whether the locked-shell
tease converts, not whether anyone wants this. The first 10 real users will
teach more than the next 10 build orders.
9.5 No distribution — the biggest non-code gap
There is no acquisition path at all. The newsletter send is unscheduled, share
cards are unbuilt (blocked on the resolution tail), social proof has no fuel
(needs a real record), partner/affiliate links are all enabled:false. A product
nobody sees cannot be validated regardless of how good the model gets. This
appears in no phase above and belongs on the board as its own track.
9.6 The read isn't actionable at the last mile
Push-to-book is a teaser — no affiliate is live, so a user who trusts a read still leaves to place it manually. Bankroll guidance (Kelly) is Desk-gated. The gap between "here's a good read" and "I placed it" is unclosed.
9.7 Operational fragility
Single-box, single Redis (persistence is a Coolify setting, not app-controlled), one cron. Settlement silently covers 2 sports. Three credentials remain flagged for rotation, including a Stripe live key that transited a chat transcript. No staging environment — every verification this session ran against prod.
THE HONEST SUMMARY
Built well: the truth infrastructure. Honest empty states, refusal paths, n-gates, the append-only ledger, the settled/live gate, the atomic founder cap. This codebase does not lie about what it knows — that is rare and it is real.
Not yet true: that the model beats the market. Not disproven either — unmeasured at adequate n, on one sport, with a projection whose best layers aren't connected.
So the finish line is not 23 orders. It is 23 orders plus a verdict from accrued data that we cannot rush — and the discipline to report that verdict honestly if it says the edge isn't there. The plan above builds the machine. Only time and honest measurement decide whether the machine is right.