Files
vyndr/specs/MASTER-PLAN.md
T
builtbykev c98338ef23 plan: add §10 — aggregator + paid-model gaps, and the one root cause behind both
Answers "what makes this the top product, not just a finished one."

THE REFRAME: the aggregator gap and the model gap are the SAME gap in two places.
Our "market" is often ONE book — MLB props are 73% single-book, and
proplineAdapter sends only {apiKey, markets} with NO regions/bookmakers param
(:152), so we take PropLine's default response. That single fact causes four
problems we had been treating as unrelated: no line shopping (the category's #1
free hook), a fair_prob_lock that is a de-vigged single soft book rather than a
consensus (the bent ruler the model is judged against), weak CLV (cannot measure
beat-the-close against one book), and no steam/disagreement detection (needs >=2
books to exist).

So the highest-leverage unblocked action in the whole plan is a cheap API test:
does PropLine return more books with a regions/bookmakers param on our tier? One
request, and if it works it upgrades the free product, the model's denominator and
the CLV instrument simultaneously.

Aggregator gaps catalogued: book breadth, true consensus, historical odds archive
(started — closing_captures 844k rows, lock_lines new, but in-grade history capped
at 24 points, so no full open->close series), market breadth (11 live vs the
category's 50+), ingested alt-line ladders, injury/lineup wire, player news.

Paid-model gaps catalogued: distribution instead of a point (distribution.js
already computes survival probabilities and rungs but is proj-v1.1, ledger-only
and lost to the champion); opportunity/playing-time projected FIRST with its own
uncertainty (the single biggest available modelling gain); per-stat models instead
of one additive index; matchup granularity that actually reaches the grade;
applied calibration; a backtest harness (blocked by the archive gap — you cannot
backtest a price you never stored); CLV as north star.

THE PATTERN: almost every model capability is ALREADY BUILT AND DISCONNECTED.
VYNDR does not have a building problem, it has a connection-and-proof problem plus
one genuine ingestion gap that starves both halves. The expensive part is largely
done, but no new feature fixes it.

Ordering principle recorded: get MLB genuinely good BEFORE replicating across six
sports — a copied-six-times thin model is six times the maintenance for the same
absent edge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
2026-07-31 22:35:16 -04:00

345 lines
19 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# VYNDR — MASTER PLAN
**Single source of truth. Sessions EXECUTE against this and UPDATE it in place.**
Created 2026-07-31 by consolidation. Nothing built in this pass.
> **HOW TO USE:** this supersedes ad-hoc re-derivation. Before any order, read the
> phase you're in. After any order, tick the item and add one line. **Do not
> re-audit anything marked KNOWN** — that redundancy is what this document exists
> to kill.
---
## 0. VERIFICATION LEDGER (what was re-checked in this pass)
**NOTHING was re-verified. No query was run.** Everything required is already
captured in 22 artifacts produced this session plus the canonical board. Per the
order's own clause — *"If everything needed is already in the artifacts, say so and
skip verification"* — this is that case.
**Taken as KNOWN (source in brackets):**
- Model architecture, all layers [`model-architecture-recovery-map.md`]
- Grade↔outcome correlations, collapse cost [`full-output-grade-mapping.md`]
- p_win calibration + holdout verdicts [`grade-diagnostic-t0.md`, `pwin-recalibration-holdout.md`]
- Market-relative edge inversion [`grade-fix-part1-investigation.md`]
- Design implemented-vs-designed, 61 items [`design-vs-build-gap-audit.md`]
- Surface states, orphans, waves [`pre-audit-status-pull.md`, `incomplete-surface-triage.md`]
- Resolution pipeline true state [`resolution-and-clv-investigation.md`, `wave3-*.md`]
- Tier/monetization + founder mechanism [`tier-structure-pull.md`, `tier-redesign-spec.md`, `build2-review-zero-report.md`]
- Sport-boundary cost [`model-architecture-recovery-map.md` §Phase 2]
**Genuinely OPEN → carried as explicit unknowns (not verified because they need a
build or a decision, not a query):** sport order (Kev's call, §A2); board-reasoning
gating (a/b/c, §C2); CLV flag decision (§D3); the ~70 undefined team colours (§B).
---
## A. PER-SPORT MODELS — Phillips doctrine: each sport is its OWN model
### A1. MLB layer stack (the reference module), bottom-up
| # | layer | state | note |
|---|---|---|---|
| 1 | Data/feeds | **BUILT** | statsapi free+unlimited; statcast; park; weather; probables. Settles end-to-end. |
| 2 | Similarity (comparable instances) | **BUILT · NOT WIRED · NOT DEPLOYED** | `python/utils/similarity.py`; grade path skips to season/recent averages |
| 3 | Archetypes (batter + pitcher) | **BUILT · display only** | classifies + renders; **does NOT feed the grade** |
| 4 | Variable weights | **PARTIAL/ASSUMED** | engine1's flat ±1.0/±0.5 deltas are hand-set, never fitted |
| 5 | Conditions (park/weather/platoon/arsenal) | **BUILT · CHALLENGER ONLY** | rides `env_*`/`challenger_*`; "measured, never served" |
| 6 | Bayesian inference | **BUILT · NOT WIRED · NOT DEPLOYED** | `python/utils/bayesian.py`, 320 ln, genuinely sport-agnostic math |
| 7 | Calibration | **MEASURED, NOT APPLIED** | isotonic qualifies on holdout (rel .1038→.0939, res .139→.123, n=119) |
| 8 | Grade ladder | **BROKEN** | letter is a factor index, r≈0.005, **inverted** (B 52.4% < C 56.9%) |
**The through-line:** layers 2, 3, 5, 6 are built and *not connected*; layer 8 is
connected and *meaningless*. MLB's fix is connection, not construction.
### A2. Sport order (OPEN — Kev decides)
Proposed by readiness × clock: **1) MLB** (reference, only qualifying model) →
**2) CFB** (has a <30-day clock; soft-market thesis) → **3) NFL****4) NBA**
**5) CBB****6) WNBA re-attempt** (abstains today) → **7) soccer** (quota-blocked).
Each gets the same 8-layer template. **No sport is abandoned — abstention is a
state, not a verdict.**
---
## B. DESIGN IMPLEMENTATION — 61 items catalogued
BUILT-TO-SPEC 20 · DRIFTED 7 · PARTIAL 16 · ABSENT 18. Ordered wire-in:
**D1-A done** (6 combat glyphs, boundary-blue completed, reaction primitives, READ-FAB).
**D1-finish done** (rationale/reveal/chips modules — *built, NOT mounted*).
Remaining: **D1-close** (mount — blocked on a ROW-GRAMMAR slot amendment + §C2) ·
**D1-B** (45 unwired glyphs + the 41-vs-74 archetype scope call) · S2 primitive set
(movement strip, crown, disagreement axis, SPLIT) · S3 article media · The Report
email + archive · Offseason artboards · **team colours: only 10 of ~80 defined —
the rest render honest-neutral until a real source exists.**
---
## C. SURFACES
**C1 — done:** Wave 1 wiring · `/compare` · `/record` · Build-1 gate.
**C2 — OPEN DECISION (blocks D1-close):** board reasoning is served ungated while
`tiers.js` declares `reasoning_visible:false`. Options (a) gate it, (b) accept as
free funnel, (c) leave unrendered.
**C3 — remaining:** `/record` **has no nav link** (the surface that justifies the
price) · `/notifications` · Offseason hub · `/system` · S3 media · `/soccer`
(quota) · share cards (blocked by D).
---
## D. RESOLUTION PIPELINE → USER OUTPUT
**KNOWN and load-bearing: settlement WORKS** (scheduler → `settleAllOutcomes` +
`settleAllLedgers`, 937+ settled, growing daily). **What is unreachable is the
user-output TAIL:** `/api/grading/resolve` has no caller, and its fanout holds
webPush/Telegram/Discord but **no share-card step and no recap**.
**D1** wire a trigger (or move the fanout into the settle pass) · **D2** share-card
generation + `/notifications` consent + result posts + recap · **D3 CLV flag
decision** — `clvCaptureReliable()` is *one env var*, and the pre-registered rule
stands: flip only if close_moved is a clear majority AND coverage is representative.
**🔴 Never wire `/api/grading/resolve` as a second settlement path — it double-counts.**
---
## E. SPORT BOUNDARY
Adding a sport is a **~10-file core edit** with four silent-failure modes
(MARKET_MAP → zero props; three stat whitelists → silent 400s; missing projection →
universal refusal; no settled feed → grades forever). **Collapse to a registry** so
a sport is a module. **Blocks all of A2 after MLB.**
---
## F. CHROME AUDIT — 11 items, 4 need a Desk session. **Runs when surfaces are stable, not before.**
---
# THE PHASES — 7 phases, ~18 orders
| phase | orders | contents | blocks |
|---|---|---|---|
| **1. MLB model truth** | 4 | promote isotonic p_win (MLB only, WNBA abstains) · rebuild the ladder on calibrated p_win · re-adjudicate (ROI-by-grade, skew, proj-v1.1, C1 floor) · connect layers 2/3/5/6 | everything model-shaped |
| **2. Resolution tail** | 3 | trigger · share cards + notifications + posts + recap · CLV flag decision | share cards, social proof |
| **3. Surfaces + design lane** *(parallel with 1-2)* | 4 | C2 decision → D1-close mount · `/record` nav + remaining surfaces · D1-B glyphs/archetypes · S2 primitives | Chrome audit |
| **4. Sport boundary** | 2 | registry collapse · MLB re-expressed as the first module | all further sports |
| **5. Sport rollout** | 1 per sport | CFB → NFL → NBA → CBB → WNBA retry → soccer, each on the 8-layer template | — |
| **6. Monetization finish** | 2 | Stripe Phase-B live proof on first real signup · founder launch to the 3 existing users | — |
| **7. Chrome audit + hardening** | 2 | the 11-item visual sweep · credential rotation + migration-drift reconciliation | ship |
**Phases 14 and 67 = ~17 orders. Phase 5 = 1 order per sport (6 listed).**
**Total ≈ 23 orders to the end state**, of which **~11 are unblocked today**.
---
# DEFINITION OF DONE
**VYNDR is complete when:**
1. **MLB layers 18 are BUILT AND CONNECTED** — similarity, archetypes, fitted
weights, conditions and Bayesian all feed the grade; calibration applied; the
ladder monotone (A>B>C, no inversion) and proven on held-out data.
2. **Every listed sport is finished on the same 8-layer template**, or explicitly
abstaining with its reason recorded — never silently absent.
3. **Design fully implemented** — all 61 catalogued items BUILT-TO-SPEC.
4. **Every surface built, reachable and honest** — no orphans, no live-but-not-honest
surface, no dead component.
5. **Resolution pipeline live end-to-end** — settle → share card / notification /
post / recap, firing on a real settlement.
6. **Sport boundary is a registry** — a new sport is a module, not a core edit.
7. **Chrome audit passed**, logged-out and entitled.
8. **The record is publishable on its own terms** — CLV either trustworthy-and-
representative or honestly absent; no claim outruns its evidence.
**The remaining work is finite and countable: ~23 orders across 7 phases.**
---
## STANDING LAWS (carried into every order)
Truth Law — absent beats wrong, no fabrication up or down · per-sport models, never
a global engine · lookahead guard (lock-time fields only) · overfitting guard (fit
one split, prove another) · never mint A's without new information · aggregate proof
is free, itemized judgment is paid · one canonical founder flag · atomicity by unique
index, never a count · cache-bust every post-deploy check · verify-after-write.
---
# 9. WHAT'S ACTUALLY MISSING FOR THIS TO WORK AS A PRODUCT
*The phases above say what is UNBUILT. This says what is missing for VYNDR to
genuinely do what it claims. Some of it is not a build, and one of it is not
fixable by us at all. Written plainly because a plan that only counts code is the
comfortable version.*
## 9.1 🔴 THE CENTRAL ONE: there is no demonstrated edge yet
Every edge measurement this session came back **null, negative, or unproven**:
| measurement | result |
|---|---|
| served grade → outcome | **r ≈ 0.005**, and **inverted** (B 52.4% < C 56.9%) |
| p_win fair_prob (3 formulations) | **negative in all three, both sports, both splits** |
| p_win alone, MLB, holdout | +0.165, **p ≈ 0.07 — not significant** |
| p_win alone, WNBA | **negative** — abstains |
| CLV / beat-close | **null by guard** — instrument not trustworthy |
| ROI by grade | likely an artifact of a meaningless letter |
**The product's core claim — "our read is better than the market" — is not
currently supported by our own data.** Everything else in this plan is
scaffolding around that. Building all 23 orders and *not* closing this leaves a
beautifully-built product that doesn't do the one thing it sells.
**What closes it:** not code. **Sample and honest iteration.** The instrument
fields are ~10 days old (442 rows). At ~90 decided MLB rows/week, a defensible
verdict is **610 weeks out**. That clock cannot be shortened by engineering, and
any attempt to shorten it is the fabrication this whole session has been removing.
## 9.2 The projection — the actual engine — is thin and unvalidated
The grade's only real inputs today are **l5/l20 averages, an opponent rank, rest
and usage**. Similarity, archetypes, park/weather/platoon and the Bayesian layer
are all built and **not connected**. So VYNDR is currently a recent-form average
wearing an intelligence system's clothes. Phase 1 connects them — **but connecting
them is a hypothesis, not a guarantee.** They must each prove out on held-out data
or be left disconnected honestly.
## 9.3 A one-sport product marketed as multi-sport
MLB is the only qualifying model. WNBA abstains on its own data. NBA and soccer
**don't even settle** — they grade into a void. Until Phase 5, the honest framing
is *"an MLB product with other sports in development."* The site should not imply
otherwise.
## 9.4 No customers, therefore no feedback loop
**3 users, 0 paid.** The founder mechanism is built and race-proven, `/record`
exists, the gate works — and **none of it has met a real user.** Nothing here is
validated by usage: not the price, not the tier line, not whether the locked-shell
tease converts, not whether anyone wants this. **The first 10 real users will
teach more than the next 10 build orders.**
## 9.5 No distribution — the biggest non-code gap
There is no acquisition path at all. The newsletter send is unscheduled, share
cards are unbuilt (blocked on the resolution tail), social proof has no fuel
(needs a real record), partner/affiliate links are all `enabled:false`. **A product
nobody sees cannot be validated regardless of how good the model gets.** This
appears in no phase above and belongs on the board as its own track.
## 9.6 The read isn't actionable at the last mile
Push-to-book is a **teaser** — no affiliate is live, so a user who trusts a read
still leaves to place it manually. Bankroll guidance (Kelly) is Desk-gated. The
gap between *"here's a good read"* and *"I placed it"* is unclosed.
## 9.7 Operational fragility
Single-box, single Redis (persistence is a Coolify setting, not app-controlled),
one cron. Settlement silently covers 2 sports. **Three credentials remain flagged
for rotation, including a Stripe live key that transited a chat transcript.** No
staging environment — every verification this session ran against prod.
---
## THE HONEST SUMMARY
**Built well:** the truth infrastructure. Honest empty states, refusal paths,
n-gates, the append-only ledger, the settled/live gate, the atomic founder cap.
**This codebase does not lie about what it knows** — that is rare and it is real.
**Not yet true:** that the model beats the market. Not disproven either — *unmeasured
at adequate n*, on one sport, with a projection whose best layers aren't connected.
**So the finish line is not 23 orders.** It is 23 orders **plus a verdict from
accrued data that we cannot rush** — and the discipline to report that verdict
honestly if it says the edge isn't there. The plan above builds the machine. Only
time and honest measurement decide whether the machine is right.
---
# 10. TO BE A REAL AGGREGATOR *AND* A MODEL PEOPLE PAY FOR
*Kev's question: what makes this the top product, not just a finished one. The
answer that matters most: **the aggregator gap and the model gap are the SAME gap
in two places.** Fix the data breadth and both halves improve at once.*
## 10.1 The one finding that reframes everything
**Our "market" is often ONE book.** MLB props are **73% single-book** (2.18 audit).
And `proplineAdapter` sends only `{ apiKey, markets }` — **no `regions`, no
`bookmakers` param** (`:152`). We take PropLine's *default* response.
That single fact causes four separate problems we have been treating as unrelated:
1. **No line shopping** — the #1 free-tier hook in this category needs many books.
2. **`fair_prob_lock` is a de-vigged SINGLE SOFT BOOK**, not a consensus. That is
the ruler the model is judged against — a bent one. (Flagged as T1; never run.)
3. **CLV is weak** — you cannot measure "beat the close" against one book's close.
4. **No steam/disagreement detection** — needs ≥2 books to even exist.
**So the highest-leverage unblocked action in the whole plan is a cheap API test:
does PropLine return more books with a `regions`/`bookmakers` param on our tier?**
It is one request. If yes, it upgrades the free product, the model's denominator,
and the CLV instrument simultaneously.
## 10.2 What a real DATA AGGREGATOR has that we don't
| capability | ours | gap |
|---|---|---|
| **Book breadth** | 5 MLB / 2 WNBA, 73% single-book | the category runs 1020. **Root gap (10.1)** |
| **True consensus / no-vig line** | single-book de-vig | needs breadth first |
| **Historical odds archive** | **STARTED**`closing_captures` 844k rows, `lock_lines` (033) new, in-grade history capped at **24 points** | no full open→close series per prop. This is what makes CLV and backtesting real |
| **Market breadth** | 11 live markets | the category ships 50+ (alt lines, combos, innings, quarters) |
| **Alt-line ladders from books** | we *compute* a ladder; we don't *ingest* the books' | users shop rungs |
| **Injury / lineup wire** | partial (`depthChart`, confirmed-vs-projected) | no real-time news wire |
| **Player news** | `NewsWire` on Explore | not beat-level, not per-prop |
**None of this is model work. It's ingestion.** And it is the half competitors
compete on hardest, because it is visible to a free user in five seconds.
## 10.3 What a prediction model people PAY for has that we don't
1. **A distribution, not a point.** We project a point (l5/l20 average) and take
an empirical `P(over)`. Paid-tier models simulate a **full distribution per
stat** (negative-binomial / Poisson / MC). **We already have this**
`projection/distribution.js` computes real survival probabilities and a rung
ladder — but it is **proj-v1.1, ledger-only, and it lost to the champion.** The
asset exists; it is unconnected and unproven.
2. **Opportunity modelled FIRST.** In props, playing time is the dominant driver —
plate appearances, snaps, minutes, batting-order slot. We carry `ab_per_game`
and minutes as *features*, not as a **projected opportunity** with its own
uncertainty. This is the single biggest modelling upgrade available.
3. **Per-stat models.** Hits, strikeouts and total bases have different shapes.
One additive factor index across all of them is why the ladder is meaningless.
4. **Matchup granularity that actually reaches the grade.** Arsenal, handedness,
park, weather, platoon — **all built, all challenger-only, none feed the grade.**
5. **Calibrated probabilities with honest intervals.** Measured (isotonic
qualifies on MLB) — **not applied.**
6. **A backtest harness on real historical odds.** Blocked by 10.2's archive gap:
you cannot backtest a price you never stored.
7. **CLV as the north-star metric**, published honestly. Instrument built,
guard-blocked, and weak until book breadth lands.
## 10.4 The uncomfortable pattern
**Almost every model capability above is ALREADY BUILT and DISCONNECTED**:
similarity, Bayesian, archetypes, park/weather/platoon, the distribution ladder,
calibration. VYNDR does not have a *building* problem. It has a **connection and
proof** problem — plus one genuine ingestion gap (book breadth) that starves both
halves at once.
That is good news: the expensive part is largely done. But it also means **no new
feature fixes this.** Connecting the layers and proving them on held-out data is
the work.
## 10.5 If I had to order it for "top product"
1. **Book breadth test + consensus fair line** (10.1) — one API call to find out;
upgrades aggregator, model denominator and CLV together.
2. **Opportunity projection** (10.3.2) — the biggest genuine modelling gain.
3. **Per-stat distributions** — connect `distribution.js`, prove per stat.
4. **Connect the built layers** (Phase 1) — each proven on held-out or left off.
5. **Full odds archive** — store every book's open→close; unlocks backtesting.
6. **Market breadth** — 11 → 50+ markets is mostly ingestion + the 4-layer wiring.
7. **Then** the sports rollout, on a template that is actually worth replicating.
**The ordering principle:** do not replicate a thin model across six sports. Get
MLB genuinely good first — a copied-six-times thin model is six times the
maintenance for the same absent edge.
## 10.6 The honest caveat on "top product"
The category's leaders are judged on one number: **do their picks beat the closing
line, at scale, published.** We cannot claim that yet — not because the product is
unfinished, but because **we have not measured it at adequate n on a market we can
trust.** Book breadth + the odds archive + accrued settlements are what make that
claim *possible*. Everything in §10 is in service of being able to make it — or of
being able to say honestly that we can't.