e81c9b8c51
Per-book prices existed only transiently (odds cache, ~1h, raw names, grade-path
input); every grade-path persistence point collapses to one book. The
/api/books feature was built+mounted but non-functional (fed FLAT rows to a
GROUPED comparator -> always empty).
Phase 1: bookPriceStore captures per-book prices from `props` BEFORE dedupeProps,
keyed nameKey|stat, into bookprices:{sport} (SNAP_TTL) in snapshotService. Fenced:
reads props, writes its own key, read by nothing on the grade path. Grade proven
byte-identical (test + no-grade-path-reference grep test).
Phase 2: scripts/measure-book-spread.js reports same-line best-vs-worst spread
(cents + implied-prob pts), per sport, never pooled. Pre-registered crown
threshold: median >=8c OR >=2pp. Runs post-deploy on real data.
Phase 3 (backend): compareProp is honest-absent (single-book/flat -> no crown)
and the crown is gated (BOOK_CROWN_ENABLED, default OFF until Phase 2 clears).
/api/books repointed to the snapshot-locked store (fallback odds cache),
nameKey-matched; `source` field is the deploy fingerprint.
HELD unchanged: dedupeProps, snapshot dedup, selector, grade, champion,
challengers, ranking, edge_pct/ev_pct. UI routing of BookComparison + crown
treatment deferred to post-measurement (gated on Phase 2). Full suite 3834 green,
web build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1
1975 lines
130 KiB
Markdown
1975 lines
130 KiB
Markdown
# VYNDR — STATE OF THE WORLD
|
||
### As of `4f3f433` (main, DEPLOYED + fingerprinted live), 2026-07-22. This file opens every future session. **Start with the CURRENT STATUS + OPEN ITEMS block below.**
|
||
|
||
---
|
||
# 🎯 CLV REDIRECT — investigated 2026-07-26. FINDING: it was ALREADY BUILT.
|
||
*The premise "CLV is dead, redirect it to closing_captures" is STALE. Nothing needed building.*
|
||
- **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows`→`closing_captures` (append-only,
|
||
provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221`
|
||
+ internal endpoint; `ledgerService.attachClosingProb`→`closing_prob` (de-vigs both raw sides,
|
||
write-once, honest-absent `market_unavailable_reason`) via `snapshotScheduler:310`.
|
||
- **Honest-absent is already implemented + is the invalidation marker:** 651,624 capture rows
|
||
(86% `missed`, no fabricated price); ledger has **59 genuine `closing_prob`, 870
|
||
`market_unavailable_reason`**. Of 841 old-`clv` rows, **791 are already marked unavailable**;
|
||
**745 (89%) of old `clv` = exactly 0/'flat'** (the fake "never captured" signal). Old broken
|
||
`clv`/`closing_line` left in place, unread by the new instrument (but STILL consumed by ledger
|
||
ROW_COLUMNS/UI — repointing display is a FUTURE order, not done here).
|
||
- **0.4 GATE:** captures **differ from locked** (10/10 comparable rows moved) → NOT a
|
||
captures==locked timing defect. BUT the capture **86%-misses** → CLV is STARVED: only **~10-50
|
||
usable rows**, and only **10 have both `closing_prob` + `fair_prob_lock` (all WNBA, 0 MLB)**.
|
||
- **LINE-LAG HYPOTHESIS → CANNOT DETERMINE.** The 10 WNBA comparable rows move ~1pp, roughly
|
||
symmetric (mean CLV gap −0.008 over / −0.011 under). **The +4.57% MLB-C and the over/under
|
||
asymmetry are NEITHER confirmed nor refuted by CLV** — there is essentially no MLB CLV data.
|
||
- **The real open defect (separate order):** WHY 86% of captures miss (props leave the feed before
|
||
the 45-min window? matching failure?). Until that's fixed the instrument stays starved.
|
||
- Docs corrected: CLAUDE.md (WNBA settles, not MLB-only), STATE.md grade_11 (model_snapshots only).
|
||
Orphan `next start -p 3111` (6-day-old, the "1 shell running") killed.
|
||
|
||
---
|
||
# 📐 CALIBRATION DIAGNOSIS — measured 2026-07-26 (read-only; nothing changed)
|
||
*Run on BOTH populations: unselected both-sides (model_snapshots, p_win per side) vs published selected (ledger). Never pooled across sport. Full detail in chat log of that order.*
|
||
|
||
**Feasibility caveats that govern every number below:**
|
||
- p_win window ≈ **last 10 games** (`probabilityEstimator`, last-5 double-weighted) → **SE ≈ 0.16 at p=0.5** (noisy estimator).
|
||
- Continuous p_win coverage is thin+recent (revived ~07-20): **~18 MLB games, ~5 WNBA games** of joinable calibration data. WNBA calibration is **effectively anecdotal (~5 games)**; MLB is suggestive not confident (~18 games).
|
||
- Effective sample generally: MLB 31 games / 190 players, **WNBA 12 games / 77 players** — prop outcomes cluster within a game; do not treat rows as independent.
|
||
|
||
**Findings (VERIFIED where measured; sample caveats attached):**
|
||
1. **Selection bias, measured directly** = both-sides→selected Brier: **MLB 0.212→0.233, WNBA 0.258→0.278 (~+0.02 each)**. Published grade is inflated ~1 sub-tier by the side-pick: mean idx **MLB C→C+**, **WNBA C+→B−**. Honest A-or-better ≈ 0 → the "caps at ~B+/A-" story survives (optimistic if anything).
|
||
2. **MLB calibration** (both-sides): Reliability 0.010 (good), Resolution 0.047 (moderate), Uncertainty 0.249 → **discriminates + roughly calibrated**; middle deciles near-perfect. SELECTED side is **overconfident, growing with p** (+0.02 at p<.5 → **+0.19 at p≥.8**).
|
||
3. **WNBA calibration** (both-sides): Brier 0.258 **> uncertainty 0.25 = worse than always-predicting-0.5**; Resolution 0.019 (near-noise), Reliability 0.027 (poor). Realized ~flat vs predicted. **Champion does not discriminate on WNBA** — but ~5 games, so unproven.
|
||
4. **Null baselines (real prices):** **unders lose ~−14% flat both sports; overs near break-even** (MLB −2.8%, WNBA −1.4%). **Ledger strongly supports "no unders by default."** MLB model beats always-over (MLB-C +4.6% vs −2.8%); **WNBA model LOSES to always-over** (−5% vs −1.4%). Both within noise at 31/12 games.
|
||
5. **ev_pct = p_win×decimal−1 uses the overconfident p_win** → inflated; hero ranks on it → picks the most-overconfident reads. No current metric reliably predicts ROI; price-aware EV requires recalibrated p_win first.
|
||
|
||
**BUILD FORK (declared before data; sample-tempered):** MLB → *resolution good / reliability poor-when-selected* → recalibration mapping + EV ranking — **but ~18 games can't validate a holdout yet (test half 15 games)**. WNBA → *resolution poor* → champion doesn't discriminate / price-aware replacement — **but ~5 games = accrue volume before concluding.** **CLV redirect ships next regardless.** Most robust actionable finding now: **kill default unders.**
|
||
|
||
|
||
---
|
||
|
||
# 🧭 CURRENT STATUS + OPEN ITEMS — orientation block (2026-07-20, ~03:20 UTC)
|
||
*Written to orient a fresh session. Ground truth, not optimism. Everything below
|
||
was verified live unless explicitly marked UNVERIFIED.*
|
||
|
||
## What shipped tonight (all deployed + fingerprinted)
|
||
|
||
| # | Shipped | Proof |
|
||
|---|---|---|
|
||
| 1 | **Probability layer revived** — it was DEAD in prod | `p_win`/`ev_pct`/`model_odds`/`value` on **32/32** live grades (was 0/8) |
|
||
| 2 | **Value engine Arc 1** — de-vig, EV, value triplet, hero v2 | live: book −140 · fair −125 · model −109, `value:false` correctly |
|
||
| 3 | **Grade-range work** — L20 symmetry, `refreshTeamStats` wired, consistency CV guard, `confidence_basis`, dead `mlbGrader.js` removed | suite 279/3325; D emitted on real WNBA data locally |
|
||
| 4 | **Backup durable ON-BOX** | dump read back: **643 rows in dump == 643 live**, 18 MB |
|
||
| 5 | **`model_snapshots` retention LIVE** (migration 025) | **100 rows** captured over 2 cycles, incl. **36 refusals** |
|
||
| 6 | ESPN team-stats parser fixed (`buckets is not iterable`, 0/15 → 15/15) | live refresh |
|
||
|
||
### ⚠️ Two honest qualifiers on the above
|
||
- **A STILL DOES NOT EMIT IN PRODUCTION.** The 11-step grade is stored in
|
||
`model_snapshots.grade_11` ONLY (CORRECTED 2026-07-26: it is NOT in
|
||
`ledger_entries` — `_grade_11` is deleted at `gradeSlateService.js:97` before
|
||
the ledger write; sub-tier resolution on a SETTLED ledger row is recoverable
|
||
only by joining model_snapshots). A/D are *arithmetically* reachable + locked
|
||
by tests, but the
|
||
±1.0 opponent factor is still dead because `opp_rank_stat` is underivable from
|
||
ESPN (see open item 6). Live boards remain **B/C only**.
|
||
**→ The "A-RATED" marketing hold STANDS** (`AccuracyBadge` correctly falls
|
||
through to "MODEL · X% HIT"; `TopSignals` self-hides). Do not promote A-rated.
|
||
- **EV IS OVERCONFIDENT AND UNVALIDATED.** First captured values include
|
||
**+62.1% / +61% / +56.9%** — real markets don't offer that. Cause: `p_win`
|
||
clamps at `PROB_CEIL` 0.95 off ~10 games. **Hero v2 ranks on `ev_pct`, so it
|
||
picks the MOST overconfident read.** Do not surface EV as truth until
|
||
calibration runs on the harness.
|
||
|
||
## 🟢 OFF-BOX BACKUP — CLOSED. FULL ROUND TRIP PROVEN (2026-07-20 05:25 UTC)
|
||
|
||
**Root cause (Kev):** the vyndr-backup key was only in Hetzner's PROJECT ssh-key
|
||
store, never in the Storage Box's own `.ssh/authorized_keys`. Installed there →
|
||
the box now offers `publickey,password` (it previously offered an EMPTY auth
|
||
list, which is what "SSH not enabled on this box" looks like).
|
||
|
||
**Proven end-to-end:**
|
||
- `POST /api/internal/backup/run` → `ok:true`, **`offbox_ok:true`**, exit 0, 31s.
|
||
- Output confirms both env questions: **`ssh key: base64-decoded`** (the key is
|
||
correct base64, not a mangled raw paste) and destination
|
||
**`u635423@...:vyndr-backups/`** — no leading-slash bug.
|
||
- **File VERIFIED on the box, not inferred from exit 0:**
|
||
`GET /api/internal/backup/offbox` (rsync `--list-only` through the pinned host
|
||
key) returned `vyndr-20260720-051158.dump`, **833,917 bytes, 05:12:28 UTC** —
|
||
byte-identical size to the local dump, timestamp from that run.
|
||
|
||
**Hardening shipped with it:**
|
||
- **Host key STATICALLY PINNED.** `ssh-keyscan` matched the out-of-band
|
||
fingerprint `SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM`;
|
||
`scripts/storagebox_known_hosts` ships to the container and the script uses
|
||
`StrictHostKeyChecking=yes` + `UserKnownHostsFile=<pin>`. `accept-new`
|
||
(trust-on-first-use) is GONE; a missing pin file REFUSES the push. Tests
|
||
assert the weakenings never return.
|
||
- **Remote dir guaranteed** — `--mkpath`, else remote `mkdir -p`. Without it a
|
||
push into a missing parent can write the dump AS the directory name: one file,
|
||
overwritten nightly, reading as "backups exist" while retaining exactly one.
|
||
- **A failed required push now PAGES** at urgent (was "low/deferred") and the run
|
||
reports `offbox_ok:false`. Exit code still tracks ON-BOX durability so a good
|
||
on-box dump never raises a false total-failure alarm.
|
||
|
||
### ✅ ROUND TRIP CLOSED — pulled back off-box and rebuilt into a live DB (05:25 UTC)
|
||
The full insurance chain is now PROVEN, not assumed: dump → validated → pushed
|
||
off-box → verified on the box → **pulled back down from the box → restored into a
|
||
live Postgres**.
|
||
- **Pulled FROM the Storage Box** (not the local copy) with the in-session key
|
||
through the pinned host key (`StrictHostKeyChecking=yes`, never bypassed):
|
||
`vyndr-20260720-051158.dump`, 833,917 bytes — byte-identical to what was pushed.
|
||
- **Restored into scratch Postgres 17** (docker): 715 archive objects,
|
||
**42 public tables**, `ledger_entries` with all **27 columns** and real rows
|
||
(spot-checked: Kelsey Mitchell points o22.5 B −105 hit).
|
||
- **ASSERTION PASSED — `ledger_entries` restored = 645, live = 645** (target was
|
||
≥ 645). `model_snapshots` also restored 100/100, so the new retention store is
|
||
covered by backups from day one.
|
||
- Scratch DB torn down; the pulled copy deleted; **both dumps still on the box**
|
||
(nothing rotated out, cron untouched).
|
||
|
||
**🔴 OPERATIONAL GOTCHA FOUND DURING THE RESTORE — record this in any DR runbook:**
|
||
the dump is written by **pg_dump 17** (Supabase is PG 17.6), so **pg_restore 16
|
||
CANNOT read it** — it fails with `unsupported version (1.16) in file header`.
|
||
The first restore attempt failed for exactly this reason. **Restore with PG 17+
|
||
tooling** (`postgres:17-alpine` works). Restoring into vanilla Postgres also logs
|
||
`errors ignored on restore: 12` — Supabase-specific roles/extensions that do not
|
||
exist locally; harmless, the data restores completely.
|
||
|
||
## Open items — with status
|
||
|
||
| Item | Status | Note |
|
||
|---|---|---|
|
||
| **Settlement: 0 pushes / 470 settled** | 🔴 OPEN, unstarted | Implausible — hits/TB land on the number regularly. Exact-number push almost certainly mis-scored as hit or miss. Corrupts every accuracy/ROI number. |
|
||
| **~28 props/day never settle** | 🔴 OPEN, unstarted | Jul 17 MLB 86 graded/57 settled; Jul 18 103/75. Cause undiagnosed. |
|
||
| **Model-version contamination** | 🟠 PERMANENT, mitigate | `ledger_entries` mixes pre/post-2026-07-19-fix grades with no marker; eras cannot be separated retroactively. **Any backtest/accuracy claim off existing ledger history MUST treat the fix boundary as a hard cutoff.** `model_snapshots` stamps `model_version`+`code_sha` so it can't recur. |
|
||
| **A-grade unreachable in prod** | 🔴 OPEN | `opp_rank_stat` null; ESPN team endpoint has no defensive metric at all. Marketing hold stands. |
|
||
| **EV overconfident** | 🟠 OPEN | Needs calibration before it drives any surface. Hero v2 already ranks on it. |
|
||
| **`edge_pct` broken scale (U-deg pt 2)** | 🔴 OPEN | 51.5% of ledger rows exceed the sane cap; worst 620. 13 frontend surfaces render it; **nothing renders `ev_pct`**; it's the free-tier hook; it's written to the append-only `edge` column every cron. |
|
||
| **CLV broken (C4)** | 🔴 OPEN | `closing_line == locked_line` on ~95% of rows. BEAT CLOSE suppressed. **CLV ledger stays PRIVATE until backtest-proven.** |
|
||
| **Consistency CV floor** | 🟠 STOPGAP | `CONSISTENCY_MIN_MEAN=4` leaves a ±1.0 dead for MLB low-count stats. Real fix = index-of-dispersion classifier; needs a backtest first. |
|
||
|
||
### 🔑 THREE KEYS TO ROTATE (all currently valid)
|
||
1. **Storage Box password** — Kev plans a fresh reset; do it before/with the SSH fix.
|
||
2. **`VYNDR_INTERNAL_KEY`** — was pasted in plaintext into a chat transcript
|
||
(2026-07-19) and is in the local `.env`. It gates snapshot triggers, backup
|
||
runs, settle passes and the newsletter send. **Rotate: Coolify env + local
|
||
`.env` together.**
|
||
3. **GitHub PAT in the `origin` remote** — plaintext in `.git/config`, leaked
|
||
since ~Session 31. Deploys run off `gitea` so it's inert, but it is a live
|
||
credential. Revoke on GitHub; the remote can simply be deleted.
|
||
|
||
## Next queued work (in order)
|
||
|
||
1. **Backtest harness** — reads `model_snapshots`. **Needs ~2 weeks of history**;
|
||
as of now it holds ONE night. Its honest first output is *"cannot validate
|
||
yet"* and that is the expected, correct result. Build the instrument now, let
|
||
data accumulate.
|
||
2. **Settlement-correctness audit** — the two bugs above. Highest-value
|
||
correctness work: everything downstream trusts hit/miss.
|
||
3. **A-grade / opponent-strength sourcing** — MLB is **solved and verified**:
|
||
`statsapi` `teams/stats?group=pitching` returns all 30 teams free with
|
||
`era/whip/avg/slg/ops/HR9`, and it's *stat-specific*, which beats what we were
|
||
fetching. NBA/WNBA need `stats.nba.com` or derived allowed-points. Implement
|
||
behind the **source-adapter pattern** (health check that PAGES on empty —
|
||
tonight's ESPN crash was invisible for exactly that reason).
|
||
4. **Metrics engine** — archetype-baselined proprietary metrics
|
||
(`specs/model-train.md` §2C). **Gated on the harness**: every metric family
|
||
ships only if it backtests better than without it.
|
||
|
||
**Governing docs:** `specs/model-train.md` (value engine + phases + north star) ·
|
||
`specs/snapshot-retention.md` (Phase 2 report + schema) ·
|
||
`specs/audit-data/grade-collapse.md` · `specs/audit-data/gate-simulation.md`.
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 80 — S7 PRICED-LINE NUDGE COMPLETED (freshness) (2026-07-22)
|
||
|
||
## 🔴 REPORT-FIRST corrected the order's premise
|
||
Order claimed "S7 is a shell that doesn't update per selection." **Not true** —
|
||
`pricedForSelection` is a `useMemo([pricedIndex, selectedPlayer, stat])` and
|
||
`setSelectedPlayer/setStat` fire on every pick → chips already update per
|
||
selection (S79 verification stood). **The real gap was FRESHNESS**: the fetch
|
||
depended on `[sport]` ONLY, so `pricedIndex` was fetched once per sport-change,
|
||
never refreshed → a page open across a pricing-cron boundary (5 UTC hours)
|
||
surfaced hour-stale lines. That was the ONLY defect, and the only thing fixed.
|
||
|
||
## Fix (display-layer only, one file)
|
||
- `refreshPriced()` re-fetches when the held snapshot is older than
|
||
**`PRICED_STALE_MS` (30s = the /api/snapshot cache)** at the moment of use —
|
||
**on selection change AND on window focus**. Long-open page never shows a
|
||
stale line. Sport change clears first (no old-sport flash).
|
||
- **Reversible gate:** whole nudge behind `PRICED_NUDGE_ENABLED = true`. False →
|
||
empties the surfaced set → S6 link-only fallback. The instant revert lever.
|
||
|
||
## Stale-tap (unchanged, already safe)
|
||
Scan submit re-fetches the live snapshot server-side, so a stale chip tapped →
|
||
real triplet (still priced) OR honest empty (rotated away). Never lands on
|
||
nothing.
|
||
|
||
## Phase 4 — every case induced live
|
||
| Case | Result |
|
||
|---|---|
|
||
| (a) priced selection | `Sam Huff hits → [u0.5 @+105]` (exact) |
|
||
| (b) change selection | `Tsung Che Cheng hits → [o0.5 @-140]` — updates |
|
||
| (c) none / off-slate | `[]` / `[]` |
|
||
| (d) tap → triplet | `book +105 · fair +120` — real |
|
||
| (e) stale-tap (8.5) | no market → **degrades to empty** |
|
||
| (f) freshness | refresh-on-stale + focus wired, fingerprint live in chunk |
|
||
|
||
## Verdict — S7 FULLY WIRED (not a shell)
|
||
Correct per-selection · updates on change · fresh on long-open · tap reaches a
|
||
real triplet · stale-tap degrades honestly. **Scan A byte-identical**
|
||
(GradeResultCard/PriceTriplet/gradeAdapter/valueState/both scan routes/
|
||
pricedLines unchanged). Independently revertible via the flag.
|
||
|
||
## Method note
|
||
Scanner page is auth-gated → data-correctness induced via the exact functions
|
||
the `useMemo` runs + the `/api/scan` endpoint the tap hits (the load-bearing
|
||
path), not the rendered gated page. Freshness proven by code fingerprint + unit
|
||
tests asserting the exact refresh logic.
|
||
|
||
## Next open (unchanged)
|
||
edge_pct/ev_pct mislabel = **already checked S(prev), no genuine mislabel** —
|
||
one metric-CHOICE flagged (GradeResultCard "EDGE" = stat-gap %, tied to the
|
||
deferred ranking decision). · No settled `p_win` yet · 3 key rotations
|
||
(PAT stripped locally; Kev must REVOKE on GitHub + rotate internal key + Storage
|
||
Box password).
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 79 — READ-CARD NO-MARKET EMPTY STATE + PRICED-LINE SURFACER (2026-07-22)
|
||
The S78 diagnosis (join works, marketless scans rightly self-hide) → made the
|
||
absence LEGIBLE without fabricating a market.
|
||
|
||
## Phase 0 gate (all clear)
|
||
- **Design-check:** built to the triplet's own REFUSAL honesty language + the
|
||
existing `EmptyState` path-forward pattern — NOT freelanced.
|
||
- **Reachability:** scanner already fetches games/odds/search; snapshot is ONE
|
||
more public 30s-cached fetch per sport.
|
||
- **Staleness = inherent guard:** every scan re-validates the market
|
||
server-side at submit, so a surfaced line that rotates → honest empty state on
|
||
tap, never a vanishing triplet.
|
||
|
||
## What shipped
|
||
- **`web/src/lib/pricedLines.js`** (pure, tested) — `indexPricedLines` keeps
|
||
ONLY snapshot rows with a real `book_odds`, keyed by exact `nameKey|stat`.
|
||
Wrong stat → []; off-slate → []; nothing suggested/interpolated/nearest.
|
||
- **`web/src/components/vyndr/NoMarketState.tsx`** — honest empty state: no
|
||
market numbers of its own; real priced lines as one-tap chips OR a link to the
|
||
board. Path forward, not a wall.
|
||
- **`scan/page.tsx`** — snapshot fetch + "PRICED TONIGHT" help chips under the
|
||
free-typed line input (help, not restriction — any prop still scannable) +
|
||
the empty state when `book_odds==null || fair_odds==null`. Tapping a chip
|
||
pre-fills + re-scans (market re-resolved server-side).
|
||
|
||
## 🔴 REVERSIBILITY (the design constraint) — VERIFIED
|
||
`GradeResultCard.tsx`, `PriceTriplet.tsx`, `gradeAdapter.js`, `valueState.js`,
|
||
`routes/scan.js` and the Next `api/scan/route.ts` are **BYTE-IDENTICAL** (a test
|
||
asserts none reference the new UI). All new behavior is in 2 added files + one
|
||
additive scan-page block. **Scan A untouched; free-tier model-lock intact.**
|
||
|
||
## Proven live
|
||
- Surfacer on live snapshot: `Sam Huff hits` → exact match returns
|
||
`[o0.5? no — u0.5 @105]`; wrong stat → `[]`; Aaron Judge (off-slate) → `[]`.
|
||
- Scan A (Sam Huff hits 0.5 under): `book 105 · fair 120` → triplet renders,
|
||
empty state hidden.
|
||
- Scan B (Sam Huff hits **3.5** over): grade D, no market → **empty state shows**
|
||
+ surfacer offers the real o0.5 line.
|
||
|
||
## Still open (unchanged; NOT this order)
|
||
- No settled `p_win` yet → adjuster measurement not begun.
|
||
- env_multiplier on a ledger row pending next MLB lock.
|
||
- mappingHarness nomination flow · A/D emission · EV overconfidence ·
|
||
edge_pct surfaces · 3 key rotations (PAT stripped locally S-runbook; Kev must
|
||
still REVOKE on GitHub + rotate internal key + Storage Box password).
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 77 — DORMANT ADJUSTERS WIRED LIVE (2026-07-21)
|
||
The four-order plumbing gap is CLOSED. park + weather + platoon now reach live
|
||
grades. Pure input-wiring — **the 5 adjuster files are byte-identical** (verified
|
||
`git diff --quiet`).
|
||
|
||
## Phase 0 gate — all three inputs AVAILABLE at build
|
||
| Input | Availability | Source |
|
||
|---|---|---|
|
||
| Venue | **ALWAYS** 15/15 | schedule game object |
|
||
| First-pitch | **ALWAYS** | schedule `gameTime` |
|
||
| Opposing-SP hand | **declared 30/30 this afternoon** (morning locks precede it → platoon absent, correct) | probable `pitcherId` → statsapi hand |
|
||
Batter hand (`statcast.bats`) + MLBAM id were already on each grade.
|
||
|
||
## What was built
|
||
**`src/services/environmentContext.js`** (the wiring; adjusters untouched).
|
||
Fetches ONCE per snapshot: schedule, probable pitchers, ONE batched handedness
|
||
call, ONE Open-Meteo forecast per home park, batter splits per hitter.
|
||
`PARK_COORDS` (30 parks) lives here as public geometry (like the dome list), NOT
|
||
in an adjuster. Best-effort throughout — any failure → that prop archetype-only,
|
||
never a pipeline break.
|
||
**`attachChallenger` is now async** + takes `contextFor(grade)` → `{environment,
|
||
matchup}` → `adjust()` (which already accepted both).
|
||
|
||
## Independent attribution (each of 4 measurable separately)
|
||
Ledger columns `env_multiplier / env_park_base / env_weather_mod /
|
||
env_weather_state` + `challenger_adjustments` jsonb keeping archetype /
|
||
environment / matchup as SEPARATE entries. Not one blended delta.
|
||
|
||
## Combined bound (tested)
|
||
Coors slugger + wind-out + favourable platoon, all at once → **δ < 0.12**. Every
|
||
layer capped, total nudge clamped. Stacks lean, never re-forecast.
|
||
|
||
## ✅ PASS CONDITION MET — env_multiplier non-null (was 0 for 4 orders)
|
||
Proven END-TO-END on the live slate (buildContext + adjust against real feeds):
|
||
`venues_with_weather 15 · opp_declared 30 · pitchers_with_hand 30`.
|
||
| Park | env | compose |
|
||
|---|---|---|
|
||
| PHI | **1.179** | park 1.09 × wx 1.082 → **δ +0.041** |
|
||
| COL/ATL | >1 | hitter parks |
|
||
| CLE 0.929 · KC 0.937 | <1 | pitcher parks |
|
||
| **TOR/MIL** | park only | `wx 1 [dome_na]` — weather stands down, **park still applies** |
|
||
|
||
## 🔴 THE ONE HONEST GAP — not yet on a LEDGER ROW
|
||
Immutability refused to rewrite today's already-locked props (`ignoreDuplicates`),
|
||
so re-inducing wrote 0 new rows and **`env_multiplier` is still 0 on ledger rows**.
|
||
The wiring is PROVEN to produce the value (above); it lands on a row at the **next
|
||
NEW MLB lock**. Verify then:
|
||
```sql
|
||
select count(*) filter (where env_multiplier is not null) with_env,
|
||
count(*) filter (where challenger_adjustments::text like '%matchup%') with_platoon
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
## Point-in-time / no lookahead
|
||
Weather = FORECAST for first pitch fetched at projection time. Platoon = split
|
||
entering the game. Neither reads settle-time.
|
||
|
||
## Still open (unchanged, NOT this order)
|
||
- **No settled row carries `p_win`** — all adjusters still unmeasurable until
|
||
tonight's locks settle.
|
||
- Read-card triplet gap · mappingHarness nomination flow · A/D emission ·
|
||
EV overconfidence · zero-push bug · edge_pct surfaces · 3 key rotations.
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 76 — PLATOON SPLITS (regressed) (2026-07-21)
|
||
**Built + tested + deployed. Not reaching live props** (same plumbing gap).
|
||
|
||
## Gates — both clear; one was ALREADY closed
|
||
1. **Splits** = statsapi pull, 1 call/hitter:
|
||
`people/{id}/stats?stats=statSplits&sitCodes=vl,vr&group=hitting`.
|
||
2. **🔴 CORRECTION — batter handedness is NOT pending.** S69's Tier-A join did
|
||
it: `statcast_aggregates.bats` = **604/604** (210 L · 327 R · 67 S). The S69
|
||
STATE note said pending; the data says done.
|
||
3. **Point-in-time holds** if fetched before first pitch (a split queried this
|
||
afternoon can't contain tonight). A HISTORICAL backtest would use
|
||
season-final and leak → **clean measurement is forward-accruing.**
|
||
|
||
## 🔑 THE SPINE — regression IS the feature
|
||
`regressed = (PA·observed + K·prior) / (PA + K)`, **K = 600 PA**, prior = the
|
||
hitter's **OWN** blended rate (not the league's). The question is *"is he
|
||
DIFFERENT vs this hand than he normally is"* → his own line is the null, so a
|
||
hitter with no evidence gets **nothing**.
|
||
Weights: `30 PA → 4.8% · 130 → 17% · 400 → 40% · 1000 → 63%`.
|
||
|
||
### The make-or-break test (both halves, tested)
|
||
| Same .310 vs LHP | weight | multiplier | δ |
|
||
|---|---|---|---|
|
||
| **30 PA** | 4.8% | 1.011 | **0.003** |
|
||
| **400 PA** | 40% | 1.096 | **0.023** |
|
||
**8× apart.** A test asserts the ratio stays **> 5** — if the regression ever
|
||
breaks, the suite fails instead of projections quietly drifting onto noise.
|
||
|
||
## Real data — most hitters get NOTHING, and that's correct
|
||
| Player | vs L | vs R | vs LHP mult |
|
||
|---|---|---|---|
|
||
| Josh Bell (S) | .259 / 126 PA | .248 / 265 PA | **1.005** (17% wt) |
|
||
| Aaron Judge (R) | .246 / 76 PA | .248 / 185 PA | **0.999** (11% wt) |
|
||
Bell *looks* like a platoon splitter until the sample speaks. Neither is
|
||
material. **Expect most hitters to get ~nothing — honest output, not failure.**
|
||
|
||
## Honest-absent — 5 distinct routes, all exactly 1.0
|
||
`batter_hand_absent` · `pitcher_hand_absent` · `splits_absent` ·
|
||
`not_applicable` (stat) · missing SIDE → falls back to prior (weight 0), never 0.
|
||
|
||
## INDEPENDENT of the environment — deliberate
|
||
Park+weather compose into ONE coefficient (both describe the stadium). Platoon
|
||
rides its **own** `matchup` slot with its own label — entangling them would make
|
||
both **harder to attribute** when the instrument scores them. Mirrored on under,
|
||
capped ±15%, **inverted for strikeouts** (higher K rate → higher K prop).
|
||
|
||
## 🔴 THE PLUMBING GAP — now FOUR orders old
|
||
`snapshotService` still doesn't attach **venue + first-pitch + opposing-pitcher
|
||
hand** per prop and pass `environment` / `matchup` into `attachChallenger`.
|
||
**Park, weather AND platoon all adjust NOTHING live.** Everything else is ready.
|
||
|
||
## 🔴 And still: no settled row carries `p_win`
|
||
All seven adjusters remain unmeasurable.
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
## Recommendation (repeated, now urgent)
|
||
**Stop adding adjusters.** The next order should be ONLY: (a) attach venue +
|
||
first-pitch + opposing SP hand to props, (b) pass environment/matchup through,
|
||
(c) confirm `p_win` lands on a settled row. The measurement stack is now far
|
||
deeper than the plumbing feeding it.
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 75 — WEATHER MODULATION (2026-07-21)
|
||
**Built + tested + deployed. NOT reaching live props** (same venue gap as park).
|
||
`effective_environment = park_base × weather_mod`.
|
||
|
||
## Phase 0 — both feeds free + keyless
|
||
statsapi `/venues?hydrate=location` → lat/lon for every park, ONE call.
|
||
Open-Meteo hourly (temp / wind speed / wind direction), **available hours before
|
||
first pitch** = when we project. Verified live.
|
||
|
||
## 🔑 THE SPINE — two weather values, NEVER crossed
|
||
| Value | Purpose | Lands in |
|
||
|---|---|---|
|
||
| **FORECAST** at projection time | **drives** the live adjustment **AND is what the instrument measures** | `ledger_entries.wx_forecast` + `env_*` |
|
||
| **ACTUAL** at game time | future SELF-DERIVED weather factors ONLY | `game_context.wx_*` |
|
||
Using the actual to measure tonight = **scoring ourselves on information we
|
||
didn't have**. The actual comes from Open-Meteo's **ARCHIVE** endpoint — asking
|
||
the *forecast* endpoint after the fact returns a re-forecast, not what happened.
|
||
|
||
## 🔴 WIND SIGN — the easiest error in the file
|
||
`wind_direction` is METEOROLOGICAL (where it comes FROM). Blowing OUT to centre
|
||
= arriving FROM the opposite bearing → **`(from + 180) % 360`** before comparing
|
||
to the CF bearing. **Getting this backwards inverts every wind adjustment.**
|
||
Pinned by a test on straight-out / straight-in / crosswind.
|
||
`CF_BEARING` (30 parks) is **public geometry**, same class as the dome list. A
|
||
park missing from it gets **no wind effect** (keeps temperature).
|
||
|
||
## Three honest do-nothing states — all 1.0, distinct
|
||
- **`dome_na`** — weather doesn't apply; **the PARK factor STILL does** (verified:
|
||
domed venue keeps its sub-1.0 base while weather stands down).
|
||
- **`forecast_absent`** — none available for this park/time.
|
||
- **`sub_threshold`** — real forecast below a meaningful bar (a 0.3% nudge on a
|
||
light breeze is false precision).
|
||
Plus `not_applicable` for stats weather says nothing about (strikeouts).
|
||
|
||
## Conservative + ledger-tunable
|
||
Every magnitude is an env var (`WX_WIND_PER_MPH` .006 · `WX_TEMP_PER_DEG` .003 ·
|
||
thresholds 5 mph / 8°F), total **capped ±12%**. Nothing asserted — a nominated
|
||
challenger that earns its place or is cut.
|
||
|
||
## Induced (Wrigley, CF bearing 32°)
|
||
`wind FROM 212° @15mph → 15mph straight OUT · wx 1.12 × park 1.06 = 1.187 ·
|
||
δ +0.043` · under mirrors · **pitcher HR-allowed moves WITH hitter HR** (both are
|
||
P(over) on a ball leaving the park) · wind IN → 0.955 · calm/dome/no-forecast all
|
||
→ 1.0 by three different routes, park base still applying.
|
||
|
||
## ⚠️ ORDER CORRECTION (recorded)
|
||
The order described wind-out as helping the hitter and **hurting** the pitcher's
|
||
HR-allowed as "opposite sides." In PROP terms both go the **same** way — the
|
||
HR-allowed OVER is also more likely. **The sign lives in the STAT**, as already
|
||
established for park. Implementation follows that, not the phrasing.
|
||
|
||
## 🔴 STILL OPEN (unchanged)
|
||
1. **Neither park nor weather reaches live props** — `snapshotService` still
|
||
doesn't attach venue+first-pitch per prop and pass `environment` into
|
||
`attachChallenger`. Everything else is ready; this one link is missing.
|
||
2. **No settled row carries `p_win`** — all six adjusters remain unmeasurable.
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 74 — PARK BASE (public, pluggable) + GAME-LEVEL CAPTURE (2026-07-21)
|
||
|
||
## Phase 0 — the grain is cheaper than expected
|
||
Settle sees a player's game-log line (date + teams), **never venue or totals**.
|
||
But **ONE statsapi schedule call per game DATE** returns every game that day with
|
||
venue + linescore + scoringPlays. **Not per game, not per prop.** Verified: 15
|
||
games, 1 call.
|
||
|
||
## Public base — the ingestion was ALREADY DONE
|
||
`src/data/parkFactors.js` (static FanGraphs, S15) **is** the public base.
|
||
`src/services/parkBase.js` converts its 100-index to the composable multiplier
|
||
(COL 128 → **1.28**) instead of ingesting a second copy.
|
||
|
||
**🔴 THE COMMODITY LINE IS IN CODE, not just a comment.** Every resolution
|
||
carries `provenance`:
|
||
- public → `proprietary: false`, *"Commodity: a public number. Not a VYNDR
|
||
derivation."*
|
||
- derived → `proprietary: true` — **reserved**, and only earned once it beats
|
||
the public base on the instrument.
|
||
**Never market "proprietary park factors" about the public base.**
|
||
|
||
**Honest-absent where even PUBLIC is thin** (`UNSTABLE_PUBLIC`: ATH/OAK/SAC/LV —
|
||
relocated club, temporary venue). Extend that set whenever a park opens or a
|
||
club moves; a missing entry means a confident factor for a stadium with no data.
|
||
|
||
## SOURCE-PLUGGABLE — the architectural point
|
||
`resolveParkBase()` is the ONLY accessor; public and derived return **identical
|
||
shapes**, so callers never branch on source. Self-derived swaps into the same
|
||
slot with nothing downstream rewritten. **A derived source with no factor
|
||
returns `absent` — it does NOT silently fall back to public** (that would turn a
|
||
commodity number into a proprietary claim).
|
||
|
||
## Game-level capture — LIVE
|
||
`game_context` (migration 035), **game grain, deduped on `game_id`, never copied
|
||
onto prop rows**. Named-purpose fields ONLY: venue → park factors · runs → run
|
||
environment · HR totals → HR factors. Nothing else stored.
|
||
- **Finals only** (an in-progress total is not a result).
|
||
- **No scoring plays → HR is `null`, not 0.**
|
||
- HR from scoringPlays is COMPLETE — every HR scores ≥ the batter.
|
||
- Runs nightly beside the settle pass (yesterday's date).
|
||
|
||
**Live fingerprint:** `POST /api/internal/gamectx/2026-07-20` → **15 games, 15
|
||
venues, 15 with HR totals, 0 duplicates.**
|
||
|
||
## 📅 ACCRUAL TARGET — "later" is a number
|
||
**150 home games/venue** (≈81/season) → **~1.9 seasons** before a self-derived
|
||
factor can be NOMINATED against the public base.
|
||
Live now: **15 games · 15 venues · 0 ready.** Track:
|
||
`GET /api/internal/gamectx/accrual`.
|
||
|
||
## Induced (live)
|
||
COL HR **+0.061** hitter · **+0.061** pitcher HR-allowed (same park, sign lives
|
||
in the STAT) · mirrored on under · SF negative · **TB weather_na=true with its
|
||
factor STILL applying** · ATH **absent** · strikeouts untouched.
|
||
|
||
## ⚠️ STILL OPEN (unchanged from S73)
|
||
**Park does not yet reach live props** — `snapshotService` still doesn't attach a
|
||
venue per prop and pass `environment` into `attachChallenger`. `parkBase` +
|
||
`parkMultiplier` + the challenger's `environment` slot are all ready; the
|
||
prop→venue lookup is the missing link.
|
||
**And still blocking ALL measurement: no settled row carries `p_win`.**
|
||
|
||
## Next
|
||
1. Attach venue → prop → `environment` (closes the park loop).
|
||
2. **Weather** — multiplies `park_base`; domes already `weather_na`; add
|
||
weather-actuals to `game_context` for later self-derived modulation.
|
||
3. Self-derived park factors at ~1.9 seasons of accrual.
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 73 — PARK FACTORS (derived, composable) (2026-07-21)
|
||
**Derivation + composition BUILT and proven. PIPELINE WIRING INCOMPLETE — see
|
||
the gap.** Layer 3 Step 4.
|
||
|
||
## 🔴 PHASE 0 — A STATIC PARK TABLE ALREADY EXISTED (nearly built a duplicate)
|
||
`src/data/parkFactors.js` — **static FanGraphs table, Session 15**, consumed by
|
||
`computeFeatures.js:374`. I almost created a second source of truth.
|
||
- **NEW derived service = `src/services/parkFactors.js`** (distinct path).
|
||
- **Checked:** the static table reaches `computeFeatures` but **NOT
|
||
`probabilityEstimator`** → it does **not** currently touch `p_win`.
|
||
- **This reframes the order:** if the champion ever does see park, adding it to
|
||
the challenger is DOUBLE-COUNTING — precisely what the S72 harness catches.
|
||
Park therefore ships as a **nominated challenger to be TESTED**, not an
|
||
assumed win.
|
||
|
||
## Derivation (ours, from real games)
|
||
statsapi `schedule?hydrate=linescore,venue,scoringplays` — **one call per date
|
||
range**. Every HR scores ≥ the batter, so **HR totals are fully recoverable from
|
||
scoringPlays**. Derived from **5,055 real games, 2022–2025**:
|
||
|
||
| Park | HR base | RUN base |
|
||
|---|---|---|
|
||
| Dodger Stadium | **1.106** | 1.023 |
|
||
| Coors Field | 1.029 | **1.099** |
|
||
| Oracle Park | **0.923** | 0.969 |
|
||
| PNC Park | **0.917** | 0.994 |
|
||
|
||
**18 parks present · 18 honestly absent** (below the 150-game regime floor).
|
||
*(Sample was a partial pull — month-chunked with `-31` end dates drops days — so
|
||
absences are inflated. Re-derive with proper date ranges before trusting counts.)*
|
||
|
||
## COMPOSABLE — the architectural point
|
||
Emits a **multiplier ~1.0**, never an additive nudge, because weather must
|
||
modulate it next: `effective = park_base × weather_mod × …`. Additive terms
|
||
don't compose (5% park + 8% wind = 1.05×1.08, **not** +13%). The challenger
|
||
converts via `Math.log(env)` into the same log-odds space the archetype nudges
|
||
use. A test multiplies a placeholder weather term on to prove the shape composes
|
||
**with no rearchitecting**.
|
||
|
||
## Honest states (three, distinct)
|
||
`absent` (thin → adjust nothing) · `present` (adjust) · **`weather_na`** (dome —
|
||
**the park factor STILL APPLIES**; the flag exists so next order's weather does
|
||
nothing there). **N/A ≠ absent** — conflating them either drops a valid park
|
||
factor or applies wind indoors.
|
||
|
||
## Directional by prop-owner
|
||
`home_runs` and `home_runs_allowed` both key off `hr_base` **in the same
|
||
direction** — the sign lives in the STAT, not the park.
|
||
|
||
## Structural breaks + shrinkage
|
||
A deviating season starts a new regime **only if the NEXT season confirms** (one
|
||
odd year = noise). Only post-break seasons used. Factors **regress toward
|
||
neutral by sample size**; fine conditioning (park×hand, park×batted-ball) stays
|
||
**unavailable** until its own larger floor (400 games).
|
||
|
||
## 🔴 THE GAP — park is NOT yet reaching live props
|
||
`adjust()` accepts `environment`, and `parkMultiplier` produces it — but
|
||
**nothing in `snapshotService` looks up a venue per prop and passes it**. Props
|
||
carry `game_time` (via gameBinder) but **no venue**. Remaining wire-up:
|
||
1. Attach venue to each prop (statsapi schedule already returns it).
|
||
2. Persist derived factors (a table or cached blob) + a refresh job.
|
||
3. Pass `environment` into `attachChallenger`.
|
||
**Until then park adjusts nothing live** — the derivation and composition are
|
||
proven, the plumbing is not.
|
||
|
||
## Next
|
||
1. Wire venue → prop → park lookup (the gap above).
|
||
2. **Weather modulation** — composes by multiplying `park_base`; domes already
|
||
flagged `weather_na`.
|
||
3. Still blocking all measurement: **no settled row carries `p_win` yet**.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 72 — TIER-1 MAPPINGS LIVE + TIER-2 HARNESS (2026-07-21)
|
||
|
||
## 🔴 PHASE 0 GATE — NO HISTORICAL OOS IS POSSIBLE (and why)
|
||
`statcast_aggregates` is **overwritten nightly by design** (Layer 1 = full
|
||
re-pull upsert) → season-TO-DATE only, **no point-in-time history**. Classifying
|
||
a player for a Jul-15 game with today's aggregate feeds in Jul 15–21 →
|
||
**look-ahead leakage**, making any "out-of-sample" verdict worthless.
|
||
**→ The harness reads the archetype vector RETAINED at grade time (S70
|
||
instrument) and runs FORWARD-ACCRUAL.** Reported, not worked around.
|
||
*(If point-in-time archetypes are ever wanted, Layer 1 must retain a dated
|
||
snapshot per refresh — a real schema change, not a query.)*
|
||
|
||
## Canonical names — asserted, not assumed
|
||
Every mapping references the axis keys the classifier **emits**; a test walks
|
||
both maps against `BATTER_AXES`/`PITCHER_AXES`. **A non-existent key would look
|
||
wired and never fire** — that's the failure this guards.
|
||
|
||
## TIER 1 — LIVE (tautological, directional, MLB-only)
|
||
| Axis (canonical) | Labels | Stat | Dir |
|
||
|---|---|---|---|
|
||
| `strikeout` | PUNCHOUT / WHIFF | strikeouts | ↑ |
|
||
| `ground_ball` | SINKER / SEAM | home_runs_allowed | ↓ |
|
||
| `fly_ball` | FLY BALL / ELEVATOR | home_runs_allowed | ↑ |
|
||
| `control` | SURGEON ARM / PINPOINT | walks_allowed | ↓ |
|
||
| `power` | SLUGGER / BOMBER | total_bases, home_runs | ↑ |
|
||
| `contact` | TECHNICIAN / SURGEON | hits ↑, strikeouts ↓ | |
|
||
| `patience` | GRINDER / SNIPER | walks | ↑ |
|
||
|
||
Average player untouched; UNDER mirrors OVER exactly.
|
||
**SPEED (BURNER→SB) is HONESTLY ABSENT** — SB never reached the aggregate store,
|
||
so Layer 2 shelved the axis. `BATTER_MAP.stolen_bases = {}`, not invented.
|
||
|
||
## TIER-2 HARNESS — `src/services/mappingHarness.js`
|
||
Three hurdles a plausible-but-worthless mapping fails:
|
||
1. **MARGINAL, not correlation** — scores the archetype against the
|
||
**projection's RESIDUAL** (`outcome − p_win`). If the projection already
|
||
knows it → `REJECT (redundant)`. **This is the double-counting catch.**
|
||
2. **OOS split by DATE, never random** — rows from one game share pitcher/park/
|
||
lineup; a random split leaks.
|
||
3. **Direction validated from held-out data; a contradicted sign is REJECTED,
|
||
never silently flipped** (flipping = fitting noise).
|
||
Floors: `MIN_ROWS 200 · MIN_HELD_OUT 60 · MIN_EFFECT 0.02` → else `INSUFFICIENT`.
|
||
|
||
**Lifecycle encoded: `nominated → live → claimed`.** A survivor RUNS LIVE and is
|
||
measured; **only the quantified public claim waits.** Nothing sits dark.
|
||
|
||
## Real candidate, end-to-end
|
||
`GROUND-BALL → hits_allowed` → **INSUFFICIENT — 0 / 200 settled rows**, because
|
||
no settled row carries `p_win` yet (S70 instrument starts at the next new lock).
|
||
Correct and expected verdict.
|
||
|
||
## Gotcha worth keeping
|
||
My first synthetic fixture **aliased the carrier selector against the outcome
|
||
draw** and manufactured a 0.038 effect where zero was injected. The harness
|
||
rejected it correctly — but via the SIGN reason, not redundancy, which is how I
|
||
found the fixture bug. Draw now uses a coprime modulus. *Lesson: when a harness
|
||
gives the right verdict for the wrong reason, check the fixture.*
|
||
|
||
## Still blocking everything downstream
|
||
**No settled row carries `p_win`/`p_win_challenger` yet.** Tier-1 adjusts real
|
||
projections now, but neither it nor any Tier-2 candidate can be evaluated until
|
||
volume accrues. Verify:
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger,
|
||
count(*) filter (where challenger_delta <> 0) moved
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
|
||
|
||
---
|
||
|
||
# 🟡 SESSION 71 — LAYER 3 STEP 2: CHALLENGER PROJECTION (2026-07-21)
|
||
**Code-complete + deployed + unit-proven. NOT yet observed writing a live row —
|
||
see the verification gap below.** Champion serves users UNCHANGED.
|
||
|
||
## What it is
|
||
`src/services/challengerProjection.js` — the champion's `p_win` **adjusted** by
|
||
the Layer-2 axis vector in **log-odds space** (a nudge can't cross 0/1 and means
|
||
the same at p=.5 and p=.9). Total capped at **0.45 log-odds** — a lean, not a
|
||
re-forecast. Every deviation is attributable: `challenger_adjustments` stores the
|
||
named axis + signed nudge. Only mechanically obvious stat↔axis pairs are mapped.
|
||
|
||
**Identical to champion by construction** where there's no signal — unremarkable
|
||
player, thin sample, unmapped stat, or no classification all return the champion
|
||
value byte-for-byte with a stated `reason`. The A/B therefore differs ONLY where
|
||
archetype-awareness could help or hurt.
|
||
|
||
**Isolation:** `adjust()` is pure; the champion field is read, never written; the
|
||
served snapshot payload is still `enriched` (untouched); a challenger throw is
|
||
caught so it can never break the pipeline it's measured inside. Statcast loads
|
||
ONCE per snapshot run — grade-time I/O stays at zero.
|
||
|
||
## Induced on real players (local, real data)
|
||
| Prop | Champion → Challenger | Via |
|
||
|---|---|---|
|
||
| Judge HR **over** | 0.42 → **0.447** | BOMBER +.22, WHIFF RISK −.11 |
|
||
| Judge HR **under** | 0.58 → **0.553** | exact mirror |
|
||
| Judge strikeouts | 0.55 → **0.55 (δ=0)** | WHIFF RISK +.11 / GRINDER −.11 **cancel** |
|
||
| Skubal Ks over | 0.60 → **0.702** | WHIFF, TRAPDOOR, CANNON aligned |
|
||
| Skubal hits-allowed | 0.50 → **0.392** | a K arm allows fewer hits |
|
||
| Bell / 12-PA sample | **unchanged** | `no_distinctive_axis` / `thin` |
|
||
|
||
## 🔴 VERIFICATION GAP — read before assuming this is recording
|
||
**Zero live rows carry `p_win` or `p_win_challenger`** (`latest_lock` still
|
||
`2026-07-21 02:58:45`). Three compounding reasons:
|
||
1. **Immutability refuses to backfill.** Every prop on the current slate was
|
||
locked BEFORE these columns existed; `ignoreDuplicates` correctly skips them.
|
||
**First live rows arrive on the next NEW lock** (new game date / new line).
|
||
2. **`ledgerWritten` OVER-REPORTS.** It counts `chunk.length`, not actual
|
||
inserts — 14 induced snapshots all reported "ledger 25" while inserting
|
||
**zero**. This misled my own verification; worth a one-line fix.
|
||
3. **`statcast_aggregates` is MLB-ONLY (1,354 rows; 0 wnba/nba).** I induced on
|
||
WNBA, where the challenger correctly no-ops for lack of mechanism data.
|
||
**The challenger currently applies to MLB only.**
|
||
|
||
**Verify with:**
|
||
```sql
|
||
select count(p_win) champion, count(p_win_challenger) challenger,
|
||
count(*) filter (where challenger_delta <> 0) moved
|
||
from ledger_entries where user_id is null;
|
||
```
|
||
Expect first non-zero after the next MLB slate locks.
|
||
|
||
## HELD
|
||
**Promotion** (challenger → serving) and **any superiority claim** stay held
|
||
until the ledger proves it on segmented volume. Baseline volume: blended n≈500
|
||
≈ 5 days; per-segment n≈200 ≈ 2–4 weeks.
|
||
|
||
## Next
|
||
1. Confirm first live champion+challenger rows after the next MLB lock.
|
||
2. Fix `ledgerWritten` to report real inserts.
|
||
3. Extend mechanism data beyond MLB (or accept MLB-only challenger).
|
||
4. Let volume accrue; then compare Brier/calibration segmented.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 70 — LAYER 3 STEP 0+1: THE MEASUREMENT INSTRUMENT (2026-07-21)
|
||
|
||
## 🔴 STEP 0 FINDING — WE HAVE BEEN FLYING WITHOUT AN INSTRUMENT
|
||
**"Is the projection calibrated / does it beat the market" has NEVER been
|
||
answerable.** Three separate causes:
|
||
- `model_snapshots` holds `p_win` but has **0 settled rows** (settle writes to
|
||
`ledger_entries` only) and only spans 07-19→07-20.
|
||
- `ledger_entries` settles 642 rows but carries **no probability at all**.
|
||
- `closing_captures` has the close but **no link to a result**.
|
||
|
||
The whole measurable universe was **35 rows**, recovered by a lossy in-memory
|
||
join. Harness agrees: `INSUFFICIENT_HISTORY, scored 35, joined 404`.
|
||
|
||
### The 35-row reading (NOT conclusive — reported for the record)
|
||
| | Brier ↓ | Directional |
|
||
|---|---|---|
|
||
| Our projection | **0.2682** | 18/35 |
|
||
| Market (de-vigged) | **0.2650** | 16/35 |
|
||
| Coin flip | 0.2500 | — |
|
||
|
||
**Both score worse than a coin flip; ours is marginally worse than the market.**
|
||
Mean model p **0.566** vs actual hit rate **0.486** → **systematically
|
||
overconfident ~8 pts** (same `PROB_CEIL` overconfidence flagged in EV).
|
||
At n=35 the SE is ~8.5 pts — **noise. We have no evidence the projection beats
|
||
the market, and none that it doesn't.** Every segment (max n=28) is honest-absent.
|
||
**Every settled row has archetype `(none)`** — Layer 2 postdates all of them.
|
||
|
||
## 🔴 THE PROJECTION IS NAIVE — verified, not assumed
|
||
`probabilityEstimator` = **game-log frequency over the line, 0.6/0.4 recency
|
||
blend, opponent nudge**. Zero archetype, zero mechanism data. Grep confirms.
|
||
|
||
## Step 1 — instrument wired (migration 033)
|
||
**The four-tuple now lands on ONE row** (`ledger_entries` — the row that
|
||
actually settles): `p_win` + `fair_prob_lock` + `archetype_vector` +
|
||
`projection_locked_at` at LOCK; `closing_prob` + `closing_captured_at` from
|
||
`closing_captures`; `outcome` at settle. Calibration and market-comparison
|
||
become plain SQL instead of a join that drops 90% of rows.
|
||
- **IMMUTABLE** — written once via the existing `ignoreDuplicates` upsert, never
|
||
re-derived at settle (that would measure a projection we never made).
|
||
- **VECTOR not label** — a text column cannot attribute "did archetype help?".
|
||
- **Honest-absent BOTH ways**: past game + no capture → `market_unavailable_reason`,
|
||
never an imputed line (calibration still scores); **future game → NOT marked**,
|
||
because a close can still arrive and premature absence is as dishonest as
|
||
imputation.
|
||
- **Closing coverage verified BEFORE reuse:** 83/83 = **100%** on graded props
|
||
since capture began. The 56,890 refusals are ungraded candidates, not our props.
|
||
|
||
## 🔴 TWO BUGS CAUGHT BY INDUCING
|
||
1. **Scheduler hook iterated `SPORTS`, which does not exist in that scope.**
|
||
Inside its try/catch it would throw `ReferenceError` every tick and silently
|
||
never run — **the instrument would have looked wired and captured nothing.**
|
||
Now `cadence.ALL_SPORTS`.
|
||
2. **`attachClosingProb` selected a `fair_prob` column `closing_captures` does
|
||
not have.** It stores raw `over_odds`/`under_odds` ON PURPOSE (S64) so the
|
||
de-vig runs later. First live run: **642 rows marked market-unavailable, 0
|
||
closes attached.** De-vig now runs via `devigTwoWay`. **Repair checked, not
|
||
assumed:** all 642 are pre-capture-era games, so the absence is TRUE; **0
|
||
capture-era rows wrongly marked.** The bug would have mis-marked every future row.
|
||
|
||
## ⚠️ THE BASELINE ACCRUES FORWARD — and does NOT backfill
|
||
Historical `p_win`/closes are gone. **Immutability means today's already-locked
|
||
rows never gain `p_win`** — the instrument starts recording at the **next NEW
|
||
lock** (verified: 0 new rows during induction because every prop was already
|
||
locked). Check with:
|
||
`select count(p_win) from ledger_entries where user_id is null;`
|
||
|
||
**Volume to a real baseline** (~100 settled/day): blended n≈500 → **~5 days**;
|
||
per-segment n≈200 → **2–4 weeks**.
|
||
|
||
## Layer 3 sequencing — RECOMMENDATION
|
||
The order proposed building archetype-awareness then proving it. **We cannot
|
||
prove anything yet**, so an adjuster shipped today is unmeasurable for weeks and
|
||
we would be back to shipping on plausibility. **Fix retention first (done), let
|
||
volume accrue, then build adjusters** — each measurable from its first settled row.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 69 — LAYER 2: MULTI-AXIS ARCHETYPES (2026-07-21)
|
||
Spec: **`docs/ARCHETYPE-AXES.md`** (every name accounted for). Layer 3 NOT built.
|
||
**Archetype still does not feed grades** — display + snapshot + future Layer 3.
|
||
|
||
## Tier-A joins — all three landed, each ONE free call
|
||
| Join | Before | After | Source |
|
||
|---|---|---|---|
|
||
| Handedness | batters **0%** | **1,354/1,354** | statsapi `/sports/1/players` (bats+throws+position) |
|
||
| True role | IP proxy (drifts) | **750/750** | statsapi stats `playerPool=ALL` (751 rows; default returns only ~57) |
|
||
| **Velo** | **53%** | **99%** (721/729) | wide `pitch-arsenals` matched BY PITCH TYPE |
|
||
**Role split (real usage):** starter 219 · reliever 365 · setup 126 · closer 40.
|
||
The velo fix is what makes CANNON/HOWITZER buildable instead of shelved.
|
||
|
||
## 🔬 AXIS INDEPENDENCE — MEASURED (467 batters / 531 pitchers)
|
||
**|r| ≥ 0.70 ⇒ one trait ⇒ COLLAPSED** (never show one trait as two archetypes):
|
||
`batter k%~whiff% +0.89` · `hard-hit%~exit velo +0.88` · `chase%~swing% +0.87` ·
|
||
`chase%~bb% −0.72` · `pitcher k%~whiff% +0.76` · `gb%~fb% −0.73`.
|
||
|
||
**The finding that mattered:** `pitcher velo ~ k% = +0.14`, `~whiff% = +0.07`,
|
||
`~gb% = +0.07`. **Velocity is NOT a proxy for missing bats** — a hard thrower
|
||
who misses no bats is a real distinct type, so it earns its own axis.
|
||
`pitcher k% ~ gb% = −0.10` → PUNCHOUT ⊥ SINKER, which IS the multi-axis thesis.
|
||
|
||
## Cut-lines
|
||
Real **p75 (distinctive) / p90 (elite)**, **per role where tails differ** even
|
||
when medians agree — reliever GB% p90 **54.1** vs starter **48.9**, both median
|
||
42.5. Floors PA≥50 / IP≥10.
|
||
|
||
## 🔴 THE FLEX FALLBACK IS DELETED (`archetypeService.js`)
|
||
`classify()` returned FLEX (mlb) / SHIELD (wnba) / CONNECTOR (nba) **at weight
|
||
1.0** when nothing scored — "could not classify" rendered as a confident
|
||
classification **with descriptive education copy**. 8 of 18 MLB players carried
|
||
it, and FLEX could never be *earned* (its only input, `utility`, had zero
|
||
writers). Every sport now does what MMA already did: **unclassified = absent.**
|
||
**Live fingerprint:** Alan Roden → `primary: null, blend: []`.
|
||
|
||
## Induced on real players
|
||
| Player | Blend |
|
||
|---|---|
|
||
| Skubal | STARTER · L · **WHIFF + SEAM + PINPOINT** (all elite) |
|
||
| Judge | **BOMBER + GRINDER + WHIFF RISK** (elite power, patient, whiffs) |
|
||
| Kwan | **SURGEON + SNIPER + SLASH**, power **absent** (0.4 barrel ≠ "low power") |
|
||
| **Josh Bell** (was DRIVER) | **empty** — "No standout profile" |
|
||
| **Alan Roden** (was FLEX @1.0, 21 PA) | **all axes absent** — "Not enough PA yet" |
|
||
| velo-less SP | keeps 3 axes; velocity+slot absent |
|
||
|
||
**`absent` (no data) is distinct from a null vector entry (measured, ordinary).**
|
||
|
||
## Gotchas
|
||
- `classify(sport, null)` **threw** — an explicit null does NOT trigger a default
|
||
parameter and every scorer dereferences its arg. Guarded in `classify`.
|
||
- Three tests asserted the fallback; updated to assert absence. One over-broad
|
||
edit briefly broke a test where a player has REAL stats and legitimately
|
||
classifies (Riley → BOMBER) — removing the FALLBACK is not removing
|
||
classification.
|
||
- `TeamHub.tsx` `|| 'zzz'` is a SORT key, not a display default — sorts
|
||
unclassified last. Correct; leave it.
|
||
|
||
## Layer 3 (HELD) reads the FULL vector
|
||
`classifyPlayer()` stores every axis; only the top ≤3 surface. Projection +
|
||
adjusters (park, weather, umpire, platoon) are the next arc, not built here.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 68 — LAYER 1: STATCAST MECHANISM DATA (2026-07-21)
|
||
Ingest + stay-current. **Layers 2 (archetypes) and 3 (projection) NOT built.**
|
||
Blueprint: **`docs/MECHANISM-DATA.md`** — the pattern every sport inherits.
|
||
|
||
## Phase 0 gate — both match rates 100%
|
||
- **Batters 40/40 · PITCHERS 66/66** (real rosters: CLE, DET, MIN, NYY, LAD),
|
||
joined by **MLBAM id** against the 713-pitcher Savant feed. Zero honest-absent
|
||
on identity — the join is an integer both systems use natively, and the
|
||
snapshot pipeline **already stores it** (`playerId`) per graded row.
|
||
- **pybaseball deliberately NOT used** — it is an MIT wrapper over the same CSV
|
||
URLs; adding it reintroduces a Python runtime in a stack where the existing
|
||
Python service is already offline. Direct axios + `savantAdapter`'s parser.
|
||
|
||
## What landed
|
||
`statcastAdapter` (6 free Savant feeds) → `statcastAggregateService` →
|
||
`statcast_aggregates` (migrations 030 + 031). **Live in prod:**
|
||
|
||
| | |
|
||
|---|---|
|
||
| Rows | **1,354** (604 batters · 750 pitchers) |
|
||
| Distinct players | 1,312 → **42 two-way players carry BOTH profiles** |
|
||
| Join rate | **1,354 / 1,354 · 0 unjoined** |
|
||
| Handedness | 677 pitchers (movement feed) |
|
||
| Sufficient / thin | 998 / 356 (PA≥50, IP≥10) |
|
||
| Arsenals | avg **4.4 pitches**, max 8, 679 with ≥3 |
|
||
| Table size | **5.3 MB** (vs 500 MB plan ceiling) |
|
||
| Pull time | ~5 s, all six feeds |
|
||
|
||
**Backfill and refresh are the SAME call** — full re-pull upserted on
|
||
`(sport, season, source_id, role)`. Idempotent + self-healing: **5 induced runs
|
||
→ still exactly 1,354 rows.** No incremental who-played bookkeeping to drift.
|
||
|
||
## 🔴 TWO REAL BUGS CAUGHT BY INDUCING, NOT BY REVIEW
|
||
1. **Two-way players broke the write.** PK `(sport, season, source_id)`
|
||
collapsed a player's batter and pitcher profiles into one key → one batch hit
|
||
the same row twice → `ON CONFLICT DO UPDATE command cannot affect row a
|
||
second time`, first chunk written, rest refused. **ROLE is in the key now**
|
||
(migration 031). Ohtani stores 419 PA / 17.2% barrel as a batter AND 85.2 IP
|
||
/ 33.7° arm angle / 52.1% GB as a pitcher — two real profiles, not one
|
||
invented player.
|
||
2. **Pitch mix was primary-pitch-only.** `pitch-movement` with an empty
|
||
`pitch_type` returns **ONE row per pitcher** (677 rows / ~700 pitchers), so a
|
||
five-pitch arsenal stored as one pitch — which would have classified every
|
||
pitcher as a one-pitch arm. Mix now comes from **`pitch-arsenal-stats`
|
||
(3,205 rows = pitcher × pitch)**; movement still supplies velo/break/
|
||
handedness folded onto the primary pitch. Skubal: 5 pitches, L, FF 96.7.
|
||
|
||
## Honesty rules (each has a test)
|
||
Absent metric → `null`, never 0 · thin sample **stored and flagged** (thin ≠
|
||
missing) · unjoined player stored with null `player_key`, joins later · **all
|
||
feeds empty → REFUSES to write** so a bad night can never blank a good table ·
|
||
`updated_at` on every row.
|
||
|
||
## Staleness — verified across simulated conditions
|
||
`0.1h → false · 26h → false · 50h → TRUE · 168h → TRUE · NEVER-BUILT → false`
|
||
(threshold 48h). **Never-built is deliberately NOT stale** — different
|
||
condition, different fix; paging on a fresh install teaches the operator to
|
||
ignore the alarm. The scheduler pages on a **failed run** AND on **silent
|
||
staleness** separately: a job that stops being scheduled never produces a
|
||
failure.
|
||
|
||
## Ops
|
||
Nightly at `STATCAST_HOUR_UTC` (default **11 UTC**, after every game is final);
|
||
kill switch `STATCAST=0`; gates `STATCAST_MIN_PA` / `MIN_IP` / `MAX_AGE_HOURS`.
|
||
**Induce:** `POST /api/internal/statcast/refresh` · **probe:**
|
||
`GET /api/internal/statcast/status`.
|
||
**⚠️ `/api/internal/*` is NOT reachable on `vyndr.app`** (no Next proxy, by
|
||
design) — use **`api.vyndr.app`**. Also: local `@supabase/supabase-js` egress
|
||
from the WSL box FAILS (`fetch failed`); the MCP tool works. Induce DB-writing
|
||
jobs on the server, not locally.
|
||
|
||
## Raw per-pitch — measured, and deliberately NOT stored
|
||
2026-07-19: **4,791 pitches · 119 cols · 3.26 MB · 680 B/row · 1.2 s** →
|
||
**~776k pitches/season = 503 MB CSV, 171 MB gzipped, ~0.85 GB in Postgres**
|
||
against a **500 MB free-plan ceiling** (DB is 26 MB today). Aggregate grain is
|
||
5.3 MB. Raw stays re-pullable from the free source if Layer 3 needs it; the
|
||
cold tier (Storage Box) is the two-tier home when it does.
|
||
|
||
## Commodity, not moat
|
||
Raw Statcast is public — anyone can pull it in a second. The edge is Layer 2
|
||
(which signals define an archetype, and where the boundaries sit), Layer 3
|
||
(projections), and the settled ledger that proves them. **Having the data is not
|
||
having an edge.**
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 67 — PRICE LAYER GATED + WIRED (2026-07-20)
|
||
|
||
## ⚠️ THE PREMISE WAS WRONG TWICE — read before trusting a "wire the fields" order
|
||
1. **There is no `/api/slate`.** The board reads `GET /api/snapshot/:sport`.
|
||
2. **The snapshot already carried the price fields.** `book_odds`, `fair_odds`,
|
||
`model_odds`, `p_win`, `ev_pct`, `value`, `takeable`, `fair_prob`,
|
||
`devig_method` were all live. Nothing needed adding.
|
||
3. **`fairLine` does not exist** (zero hits repo-wide). The Option-A
|
||
"derive from fairLine" instruction had no referent — `fair_odds` is already
|
||
a real de-vigged American price.
|
||
|
||
## Phase 0.5 gates — all PASS, verified numerically
|
||
- **Is it a price?** Yes. `fair_odds` = American odds from
|
||
`impliedProbToAmerican` inside `devigTwoWay`; `fair_prob` is the probability.
|
||
Both distinct from `line` (the stat threshold).
|
||
- **Numeric match?** **8/8 EXACT.** Recomputed `fair_odds`+`fair_prob`
|
||
independently from stored raw over/under prices; every value matched to the
|
||
integer and to 3dp.
|
||
- **Poison-independent on quarantine rows?** **YES, proven on the quarantined
|
||
cohort itself.** `devigTwoWay(over, under)` takes market prices only — no
|
||
model term is reachable — and the 8 rows recomputed above are all
|
||
`wrong_opponent_grade`. The poison is in the GRADE, not the price. So
|
||
**quarantine suppresses the MODEL leg only; the fair leg stands.**
|
||
|
||
## 🔴 THE REAL FINDING — the model price was PUBLIC
|
||
`GET /api/snapshot/:sport` is unauthenticated and was serving `model_odds`,
|
||
`p_win`, `ev_pct`, `value`, `takeable` to anonymous callers on **25/25 live
|
||
rows**. The Session-66 gate on `/api/analyze` was bypassed entirely.
|
||
**`/api/hero-prop` was a SECOND door** — it reads Redis directly via
|
||
`heroPropService`, so the first fix didn't cover it.
|
||
|
||
**`src/utils/snapshotGating.js` strips FIVE fields, not one** —
|
||
`model_odds` is the price, **`p_win` IS the price in another base**, and
|
||
**`ev_pct` is INVERTIBLE** (ev = f(p_win, book_odds); book_odds is public, so
|
||
ev hands the price over). `value`/`takeable` bound it. Market legs
|
||
(`book_odds`, `fair_odds`, `fair_prob`, `overround`, `devig_method`) stay on
|
||
**every** tier — **the fair leg is never the paywall**.
|
||
`src/utils/requestTier.js` resolves tier from a bearer token when present and
|
||
**FAILS CLOSED to `free`** on every error path, so a resolution failure can
|
||
only withhold, never leak. Responses now vary by entitlement → `/:sport` and
|
||
`/hero-prop` emit `private` Cache-Control for authenticated callers and both
|
||
Next proxies forward the bearer token.
|
||
|
||
## Read card — why it was empty, and the fix
|
||
A manual scan carries **no market** (`{player, stat, line, direction}`), so the
|
||
engine has no over/under to de-vig — `book_odds`/`fair_odds` were legitimately
|
||
absent. `lookupSnapshotPrices` in the Next `/api/scan` route recovers them from
|
||
the pre-graded snapshot (the same cache-only read it already did for locked
|
||
odds + team). Join is EXACT on **player + stat + line + side** (`fair_odds` is
|
||
side-specific) and returns nothing unless book AND fair are both present — a
|
||
user-chosen line the board never graded borrows no other line's price.
|
||
|
||
## 🔴 LIVE INDUCTION CAUGHT A SECOND CLASS OF BUG
|
||
With the leg stripped, `LiveHeroProp` forwarded everything to `PriceTriplet`
|
||
**except `model_price_locked`** → `deriveValueState` fell to the
|
||
missing-model-price branch and the landing page told every anonymous visitor
|
||
**"MODEL READ WITHHELD"** — quarantine's copy, which says our own data is
|
||
untrustworthy. Nothing was poisoned; it was the paywall.
|
||
**A paywall must never wear poison's face.** Unit tests and markup review both
|
||
passed this; only rendering the deployed page caught it.
|
||
|
||
## Live fingerprint (anonymous, cache-busted)
|
||
- Snapshot: model fields `[]`, market fields intact, `model_price_locked: true`
|
||
- Read card `/api/scan`: `book -150 · fair -126 · model null · locked true`
|
||
- Landing hero renders: **BOOK −153 · ◆FAIR −129 (amber) · MODEL [LOCK BAR]**,
|
||
verdict **"MODEL PRICE ON ANALYST"** + UNLOCK CTA
|
||
- `/dashboard`, `/ledger`, `/parlay` — 200, no triplet, unchanged. No consumer
|
||
reads `ev_pct`/`model_odds`/`p_win`, so the strip broke nothing.
|
||
|
||
## 📊 FAIR-LEG ABSENCE: 0.0%
|
||
636 graded rows: 636 with book, 636 with fair, **0 one-sided**. The hero number
|
||
is not a sometimes-number on current data.
|
||
|
||
## ⚠️ PRODUCT CALL PARKED FOR KEV
|
||
The landing hero is anonymous, so it now shows the LOCK instead of the full
|
||
triplet it showed this morning (`VALUE +21.1% VS FAIR`). That follows the
|
||
stated free-tier rule exactly, but trades a strong shop-window moment for gate
|
||
consistency. **Reversal is one line** — add `model_price: true` to the free
|
||
tier in `src/config/tiers.js`, or special-case the hero route.
|
||
|
||
## Still open
|
||
- **Quarantine flag is NOT available at snapshot build.** No service writes
|
||
`quarantine_reason` (manual SQL only); it's a retrospective judgement, so a
|
||
fresh row can never carry it. The component handles the state if it ever
|
||
arrives; wiring an always-null field was refused.
|
||
- **Slate row** still held — `lib/marketBreadth.js` computes LINES only.
|
||
- **Ledger / public profile** still out — no fair-odds column.
|
||
- Screenshots unavailable again (browser capture tool errored all session);
|
||
substituted computed-style readback per state.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 66 — PRICE-LAYER TOKENS + THE PRICE TRIPLET (2026-07-20)
|
||
Path 3: establish the token foundation, build the hero feature natively on it,
|
||
migrate nothing. Design source = `specs/design-reference/Vyndr Price
|
||
Triplet.dc.html` (ACT 01) + `HANDOFF.md`, committed in `f549422`.
|
||
|
||
## ⚠️ THE PHASE-0 FINDING — the token layer already existed
|
||
The order asked me to establish a token layer. **It was already there and
|
||
already correct.** `web/src/app/globals.css` `:root` carries HANDOFF's
|
||
surfaces / borders / text ramp / fonts / grade colours byte-for-byte (aligned
|
||
2026-07-16), and `web/src/lib/colorContract.js` already encodes green-is-edge-
|
||
only + glow-is-A-tier-only with an enforcing test. Building a second parallel
|
||
layer would have created two competing sources of truth, so this session
|
||
**EXTENDED** the existing one. If a future order says "establish tokens", check
|
||
globals.css first.
|
||
|
||
**Vehicle (confirmed against the stack, not assumed):** Tailwind v4 CSS-first
|
||
(`@import "tailwindcss"`, NO tailwind.config file) + inline `style={{}}`
|
||
components reading `var(--x)` — 2,305 inline-style sites and 1,916 `var(--…)`
|
||
reads. **CSS custom properties are the only vehicle this stack consumes
|
||
natively.** A TS token module or a Tailwind theme config would have fought it.
|
||
|
||
## What was ADDED (additive proof: 149 insertions, 0 deletions in existing files)
|
||
- **One genuinely new colour** — the priced-out blue `--priced-out #8fb2de`
|
||
(+ dim/tints/borders). It is the carve-out that lets a row say "there IS
|
||
edge" without green saying "take it".
|
||
- `--glow-a` (the A-tier glow finally tokenized — the law lived in
|
||
`gradeGlows()` with no value to hand back) and the fair-leg tints.
|
||
- **The LAWS are written INTO the token block** and asserted by
|
||
`tests/unit/priceTriplet.test.js`: green = takeable edge only · glow =
|
||
A-tier only · amber = caution + the fair leg · red = miss/negative only ·
|
||
blue = edge priced out · JetBrains Mono = all data. A test also asserts every
|
||
newly-declared name is NEW (zero collisions, zero overrides).
|
||
|
||
## The law layer — `web/src/lib/valueState.js`
|
||
ONE function (`deriveValueState`) answers "is this value?"; the component
|
||
renders what it returns and never re-derives a verdict.
|
||
- **VALUE requires `ev >= 2` AND a takeable price (−160…+200). Both.** Raw
|
||
positive EV is NOT value — +11.7% EV at −210 is a juiced price we won't call
|
||
a play. The constants MIRROR `src/config/valueEngine.js` and a test reads
|
||
BOTH files and fails on drift (the `playerName.js` pattern).
|
||
- Five states + a sixth display state: VALUE · PRICED_OUT · NO_EDGE ·
|
||
QUARANTINE · REFUSAL · NO_VERDICT_LOCKED (free tier — entitled data withheld,
|
||
which is NOT quarantine and NOT refusal).
|
||
- **Free tier gates at the WIRE**: `tiers.js` gains `model_price`,
|
||
`utils/tierGating` deletes `model_odds` and sets `model_price_locked`. Book
|
||
and fair pass through on EVERY tier — **the fair leg is never the paywall**.
|
||
|
||
## 🔴 INDUCTION CAUGHT A REAL BUG THAT REVIEW WOULD NOT HAVE
|
||
The "VS FAIR" figure compared **BOOK to fair**, so state 1 rendered
|
||
**"VALUE · −6.5% VS FAIR"** — a contradiction on screen, from code that read
|
||
fine and passed its tests. The design's own two worked examples pin the
|
||
formula: **MODEL minus FAIR in implied-probability PERCENTAGE POINTS**
|
||
(book +125 · fair +110 · model +98 → +2.9; book +118 · fair +104 · model +112
|
||
→ −1.8). `modelVsFair` now reproduces both exactly and a test locks them.
|
||
The figure renders **only when its sign agrees with the verdict** — a row that
|
||
clears the EV bar on the book price while our price sits level with fair leads
|
||
with the EV instead of a number that reads as a contradiction.
|
||
|
||
## Proof (induced in a real browser; computed styles, not markup)
|
||
| State | Rendered |
|
||
|---|---|
|
||
| VALUE | model + verdict `rgb(0,212,160)` — the ONLY green in the set |
|
||
| EDGE · NOT TAKEABLE | verdict `rgb(143,178,222)`, model leg **white, not green** |
|
||
| NO EDGE | grey verdict, stated at full voice |
|
||
| QUARANTINE | `MODEL —` withheld, **book + fair still render**, amber |
|
||
| REFUSAL | no legs, no gauge — "CAN'T PRICE THIS ONE" |
|
||
| FREE TIER | lock bar; book −120 + fair −104 still honest |
|
||
|
||
**Live landing hero (real `/api/hero-prop` data, not samples):** book −153 ·
|
||
fair −129 · model −343 · `VALUE +21.1% VS FAIR` · +28.1% EV.
|
||
**Mobile 390:** three legs hold at 117px each, right edge 378 < 390, no
|
||
horizontal overflow, fair no smaller than its neighbours.
|
||
**Post-deploy fingerprint:** tokens resolve in prod (`--priced-out: #8fb2de`,
|
||
`--glow-a`), triplet live on the landing hero.
|
||
**Additive proof:** `/dashboard` and `/ledger` return 200 with NO triplet and
|
||
were not touched — 0 deletions across every existing file.
|
||
|
||
*(Screenshots were not captured — the browser screenshot tool errored on every
|
||
attempt this session. Substituted stronger evidence: computed-style readback
|
||
per state and measured layout at a real 390px column.)*
|
||
|
||
## Scope held / deferred
|
||
- **Slate row HELD** — `lib/marketBreadth.js` computes LINES only (median book
|
||
line vs model value); it has no prices at all, so the triplet needs new
|
||
upstream fields there. Not a same-pass fit.
|
||
- **Ledger + public profile OUT** — `ledger_entries` has no fair-odds column
|
||
(`dclv_fair_lock` is 0/707). Needs a schema decision.
|
||
- **The rest of the redesign is a ranked migration backlog** — see the
|
||
Session-65 inventory. Nothing was migrated this order.
|
||
|
||
|
||
---
|
||
|
||
# 🟢 SESSION 65 — TWO TRUTH-LAW FIXES (2026-07-20, from Chrome's logged-out audit)
|
||
Both found by auditing the product as a STRANGER sees it. Neither touches the
|
||
triplet build; push scoring untouched.
|
||
|
||
## Doc drift corrected in this session
|
||
The header of this file was stale at `7ea0af2` and this file carried **no record
|
||
of the directional-CLV session** (`dcdad60` → `ca9ca34`: `src/services/
|
||
directionalClv.js`, the Analyst+Desk server-gated `ClvBadge`, and removal of the
|
||
legacy C4 CLV chip from five surfaces). `BUILD-STATE.md` is stale to April 2026 —
|
||
**`specs/STATE.md` is the authoritative state file**; treat BUILD-STATE as archive.
|
||
|
||
## FIX 1 — "CLV-VERIFIED" UN-CLAIMED ON THE PUBLIC PROFILE (copy + 2 render blocks)
|
||
`/u/[handle]` was titled **"CLV-verified record"** with the subhead **"closing-line
|
||
value included"** while **zero** closing-line value renders there. Verified live,
|
||
not assumed: `GET /api/profiles/vyndr` → `beat_close_pct: null` (gated behind
|
||
`clvCaptureReliable()` / `CLV_CAPTURE_RELIABLE`, which is unset while C4 is open).
|
||
A verification claim with nothing verified shown is an unearned claim, and it
|
||
leaked the standing CLV hold onto a surface that had never been audited.
|
||
|
||
**EIGHT instances found — two of them ONLY by the post-removal residual sweep:**
|
||
|
||
| # | Where | Was |
|
||
|---|---|---|
|
||
| 1 | `u/[handle]/page.tsx` | `<title>` / `og:title` / `twitter:title` — "CLV-verified record" |
|
||
| 2 | `u/[handle]/page.tsx` | meta + og + twitter description — "closing-line value included" |
|
||
| 3 | `u/[handle]/opengraph-image.tsx` | no-record fallback tagline |
|
||
| 4 | `u/[handle]/portrait/route.tsx` | no-record fallback tagline |
|
||
| 5 | `PublicProfile.tsx` | house subhead — "closing-line value included" |
|
||
| 6 | `PublicProfile.tsx` | `✓ CLV-VERIFIED` badge (gated, dark, one env flag from re-claiming) |
|
||
| **7** | `opengraph-image.tsx` | **eyebrow `CLV-VERIFIED RECORD · 30D`** — missed in the first report |
|
||
| **8** | `portrait/route.tsx` | **same eyebrow** — missed in the first report |
|
||
|
||
Instances 3 + 4 printed the verification claim **exactly when there is no record
|
||
at all** (the `rec == null` branch), i.e. the claim was loudest where the data was
|
||
most absent.
|
||
|
||
**Applied:** honest copy on 1–5 + 7–8 (settled record · wins and misses · nothing
|
||
curated); the badge (6a) and the `BEAT CLOSE` figure (6b) REMOVED from the public
|
||
profile, OG card and portrait card. **DISPLAY ONLY — `beat_close_pct`,
|
||
`clvCaptureReliable()`, `closing_captures` and the whole CLV data path are
|
||
untouched**, and the earned directional badge stays Analyst+Desk (never public).
|
||
The claim returns here when CLV genuinely renders here.
|
||
|
||
**Title double-append fixed and VERIFIED ON COMPOSED OUTPUT, not source:** the page
|
||
title carried its own `· VYNDR` while `layout.tsx`'s metadata template already
|
||
appends `%s · VYNDR` → live was `…@vyndr · VYNDR · VYNDR`. Built the app, served
|
||
it, and read the real HTML: `<title>Settled record — @vyndr · VYNDR</title>` —
|
||
exactly one suffix. Note `openGraph`/`twitter` titles **bypass** the template, so
|
||
they carry the brand suffix explicitly.
|
||
|
||
Both share images re-rendered locally (200 / image/png) and eyeballed: no overflow,
|
||
no clip, no residual CLV scaffold.
|
||
|
||
**POST-DEPLOY FINGERPRINT (live, logged-out, `f5156dd`):** `<title>Settled record
|
||
— @vyndr · VYNDR</title>` (one suffix), og/twitter title + description un-claimed,
|
||
**0 occurrences** of CLV-verified / BEAT CLOSE / "closing-line value included" in
|
||
the served HTML, and the live OG PNG re-rendered clean. Player pages fingerprinted
|
||
in Chrome across **MLB (Josh Bell, 1 graded prop — the exact row that produced
|
||
"74"), WNBA (A'ja Wilson, 2 props), NBA (Jayson Tatum, 0 props)**: FORM and MATCHUP
|
||
render `—` with intact layout and **zero console errors**; USAGE/REST still show
|
||
real values where they exist.
|
||
|
||
## FIX 2 — FORM WAS FABRICATED ON THE PLAYER PAGE (un-fabricate; steps 2+3 HELD)
|
||
Chrome saw Josh Bell as **FORM 74** (player page) and **Form 63** (generated read)
|
||
in one session. Traced both. **Neither feeds the grade** — proven, not assumed:
|
||
`engine1.gradeProp` runs at `analyzeViaEngine1.js:446`, `buildIntelFields` (which
|
||
computes form) decorates the already-graded object at `:468`; `engine1.js` never
|
||
reads a `form` key (only raw `l5_avg`/`l20_avg` **against the line**), and nothing
|
||
in `src/` reads `result.form`. Grade inputs are byte-identical before/after.
|
||
|
||
- **63 is real** — `70 + (l5_avg / l20_avg − 1) × 60`, per-prop, per-stat.
|
||
- **74 was manufactured** — `70 + 4 × (count of tonight's graded props)`, capped
|
||
92. **Nothing on the HTTP path ever sets `stats.form`**, so that fallback WAS
|
||
the live value. Confirmed empirically: Bell's payload carries `activeProps`
|
||
length **1** → 70 + 4 = **74**, exact.
|
||
- It **poisoned a second tile**: `MATCHUP` was `gradeFromForm(form)` — Bell's
|
||
"B" was `gradeFromForm(74)`, from the prop count, with **no opponent input on
|
||
that path at all**; the no-archetype branch hardcoded `'B'`. Both fabricated.
|
||
- **SYSTEMIC, not Bell-specific** — `buildIntel` is the unconditional path for
|
||
every player and every sport.
|
||
|
||
**Applied (absent renders absent):** FORM and MATCHUP now emit `kind:'plain'`
|
||
with `'—'` when there is no real value — `'plain'` deliberately, so the page
|
||
draws no progress bar and computes no width/colour off a null. `gradeFromForm`
|
||
deleted; the prop count is no longer passed into `buildIntel` at all.
|
||
`computeFormScore`'s hardcoded **75** (returned whenever the baseline was missing
|
||
or zero) now returns `undefined` → the card's self-hiding intel section drops the
|
||
row. Induced live across **MLB / NBA / WNBA**: all render `—` cleanly; real values
|
||
(`USAGE 3.6 AB/G`, `REST B2B`) still render — the fix is surgical to fabrications.
|
||
|
||
**HELD (needs a per-sport headline-stat design call, deliberately NOT built):** a
|
||
real player-level form metric + label disambiguation (page-level vs per-stat). It
|
||
lands as `stats.form` / `stats.matchup` and lights these tiles straight back up.
|
||
|
||
### Audit of the same file (order item 10) — reported, NOT changed
|
||
`toNum()` (`playerIntelService.js:18`) returns **0** for absent, feeding
|
||
`classifierInput`: `k_rate` (`pa > 0 ? … : 0`), `ip_per_start`, `k9`, `saves`.
|
||
These are **classifier inputs, not display** — a 0 there can misclassify an
|
||
archetype (0% K-rate reads as elite contact). Same family as the `Number(null)
|
||
=== 0` rule. `mlbLast10Rows` renders `${st.hits ?? 0}-${st.atBats ?? 0}`, which
|
||
prints a real-looking `0-0` line if a game-log row lacks its stat object.
|
||
**Ambiguous — changing classifier inputs shifts archetypes, so it needs its own
|
||
order.** Display rows elsewhere (`mlbSeasonRows`, ESPN season rows) correctly use
|
||
`?? '—'`.
|
||
|
||
|
||
---
|
||
|
||
## 🟢 MLB opp_rank_stat LIVE — computed, CONSUMED, retained (2026-07-20 16:51 UTC)
|
||
- **Consumption path verified BEFORE wiring** (a value the grader doesn't read
|
||
is a dead end): `featureCache.teamFeatures` sets `out.opp_rank_stat` →
|
||
`engine1.computeFactors` READS `features.opp_rank_stat` → fires
|
||
`weak_opponent_defense` (≥0.70) / `top_opponent_defense` (≤0.30).
|
||
- **Source:** `mlbStatsAdapter.getTeamPitchingStats` — all 30 teams in ONE free
|
||
unauthenticated statsapi call. Only fills when the ESPN path yields nothing
|
||
(always, for MLB), so WNBA is untouched.
|
||
- **CONSUMPTION PROVEN on the exact retained vector** (Josh Bell hits o0.5),
|
||
varying only `opp_rank_stat`:
|
||
|
||
| opp_rank_stat | grade | factors |
|
||
|---|---|---|
|
||
| 0.367 (as retained, mid) | **B-** | l5_hot_vs_line, l20_over_line, back_to_back |
|
||
| 0.98 (weak) | **B** | + **weak_opponent_defense** |
|
||
| 0.02 (tough) | **C** | + **top_opponent_defense**, away_vs_top5_defense |
|
||
| absent | **B-** | identical to mid — neither fires ✅ |
|
||
|
||
- **RETAINED** grade-time in `model_snapshots.features` (immutable, joinable).
|
||
|
||
### 📊 COVERAGE FIRST — then the grade read
|
||
**8 of 14 graded MLB rows carried a non-null `opp_rank_stat` (57 %); 6 null.**
|
||
The nulls are almost certainly unresolved `opponentAbbr` (`teamFeatures` returns
|
||
early without one), NOT the sample floor — but that is **undetermined without
|
||
container logs**, so it is reported as such rather than asserted.
|
||
|
||
**🔴 THE FACTOR FIRED ON ZERO ROWS THIS CYCLE.** All 8 non-null values were the
|
||
SAME mid-scale `0.367` (one opponent, neither threshold crossed:
|
||
`would_fire_weak 0`, `would_fire_tough 0`). **So this slate cannot test
|
||
A-emission at all** — the factor was present but inert. Any "A didn't emit"
|
||
reading from tonight is meaningless.
|
||
|
||
**Live MLB 11-step distribution (14 graded):** `C+ 4 · F 4 · B- 2 · C 2 · D 2`.
|
||
**No A.** Notably **F now emits (4 rows)** — the range has widened DOWNWARD, as
|
||
the restored ±1.0 makes Σδ ≤ −2.51 reachable for the first time.
|
||
|
||
**A-rated marketing stays HELD.** No thresholds changed. A genuine test needs a
|
||
slate with a weak-BAA opponent (≥0.70) so the +1.0 can actually lift a read.
|
||
|
||
### Thin-sample floor in production
|
||
No thin-sample opponent appeared on this slate, so the production null-floor path
|
||
was **not exercised live** — unit-tested only. Stated rather than implied.
|
||
|
||
## 🟢 HARNESS ARMED + CLOSING CAPTURE STARTED (2026-07-20 ~16:10 UTC)
|
||
Both verified by INDUCING the real code path — never by waiting for a slot.
|
||
|
||
**A — Backtest harness runs on OUR scheduler.** `snapshotScheduler` fires it at
|
||
`HARNESS_HOUR_UTC` (default 14), appends to `harness_results`, and
|
||
`opsWatch.harnessStaleAlarm` pages if it stops (a validator that isn't running
|
||
looks exactly like one that keeps passing). No external dependency.
|
||
**INDUCED live:** `verdict INSUFFICIENT_HISTORY, scored 35, joined 283,
|
||
snapshots 478` — the correct, passing output. `POST /api/internal/harness/run`.
|
||
|
||
**B — Closing capture (CAPTURE ONLY; no CLV metric, as ordered).**
|
||
- **C4 diagnosed:** `closing_line` is ONE overwritable field with no timestamp
|
||
and no provenance. `captureClosing` writes the current line and, on a prop
|
||
match failure, silently leaves the earlier value (= the lock) in place — so
|
||
"captured a real close" is indistinguishable from "never updated". It is
|
||
**92 % equal, not 100 %** (56 rows DID move): the defect is **provenance**.
|
||
- **`closing_captures` (migration 029)** — append-only, never overwritten, BOTH
|
||
raw side prices retained so the existing de-vig engine computes the fair
|
||
closing probability later. Sharp (`pinnacle`) vs book close kept distinct:
|
||
"beat the market" and "beat our book" are different questions.
|
||
- **JOIN drops `line` on purpose** — a close that MOVED off the graded line is
|
||
the whole point of CLV. Verified safe: all 164 identity groups have exactly
|
||
ONE line per `(sport, player_key, stat, side, game_date)`; zero ambiguity.
|
||
- **Wired into the intraday refresh** (already polls ~20 min through slate
|
||
hours) with a capture-rate alarm — a missed close is unrecoverable.
|
||
- **🔴 ODAWA IS NOT CAPTURABLE.** `odawa` exists only as a UI preference option
|
||
in onboarding/settings — no adapter, no `ALLOWED_BOOKS` entry, no feed. CLV
|
||
vs our own book cannot be measured until a real Odawa feed exists.
|
||
- **Refusals proven first (test-first):** unbound game_time, doubleheader
|
||
ambiguity, missed window, or a one-sided price → `missed_reason` and **NO
|
||
price**. Induced live: 1,076 rows correctly refused `one_sided_price`.
|
||
|
||
### ⚠️ MY OWN ERROR, CAUGHT AND REVERTED
|
||
To force a capture I widened the window to 1,200 min and **persisted 7,114
|
||
rows** — every one had a lead time of **392–602 minutes**, i.e. 6.5–10 hours
|
||
before lock. **Those were not closes.** Leaving them in a table named
|
||
`closing_captures` would have poisoned the first CLV computation with pseudo-
|
||
closes. **All 7,114 deleted; table verified back to 0.** Verification is now
|
||
done with `?dry=1`, which proves the pricing path (wnba: 486 rows priced, 0
|
||
missed) while writing nothing. The real 45-min window currently yields 0 rows
|
||
because no game is near lock — which is the correct answer, not a failure.
|
||
|
||
## 🟢 HEAL EXECUTED (Order 2, 2026-07-20 ~10:00 UTC) — record is clean
|
||
Rollback point first: `vyndr-20260720-093821.dump` (856,890 B) **verified ON the
|
||
Storage Box**, not just exit 0.
|
||
|
||
**⚠️ THE DRY-RUN OVERTURNED MY OWN DIAGNOSIS — read this before trusting the
|
||
Order 1.5 notes above.** Across every bindable row the stored `game_date`
|
||
matched a real game (**MIS-DATED: 0**). The players I cited as proof of
|
||
mis-dating were **genuine DNPs** on their true dates (Freeman 07-18;
|
||
Kwan/Hedges/Davis 07-17 — their teams played, they did not). **The original 64
|
||
`player_dnp` voids were CORRECT and my Order 1.5 revert undid correct work.**
|
||
|
||
| Axis | Result |
|
||
|---|---|
|
||
| Settlement scorer | **578/578 re-computed identical, 0 differ** — confirmed, no heal needed |
|
||
| Dates | 527 correct · **0 mis-dated** · 170 unbindable (my dry-run's parser, NOT proven damage) |
|
||
| Grade damage | **51 rows, not 303** — the rest hit the same opponent anyway (baseball SERIES effect) |
|
||
|
||
**Writes executed (each verified against the dry-run prediction):**
|
||
- **64 re-voided** (60 mlb + 4 wnba) — by the settle pass's positive-evidence
|
||
guard, not hand-picked ids. DNP voiding re-enabled deliberately because its
|
||
disable rested on a false premise.
|
||
- **25 ledger rows quarantined** `wrong_opponent_grade` (wnba 2026-07-18,
|
||
10 hit / 15 miss). They REMAIN real public settled results; only train/validate
|
||
excludes them.
|
||
- **200 `model_snapshots` rows quarantined** — the 4 pre-fix MLB cycles for
|
||
game_date 07-20 (06:00/06:02/06:03/07:25 UTC). Superseded by the correct
|
||
07:50/07:51 captures of the same props, so nothing is lost.
|
||
- **70 rows flagged** `doubleheader_game_identified`.
|
||
|
||
**TWO EXCLUSION SCOPES (migration 027) — deliberately separate:**
|
||
- `quarantine_reason` → grade untrustworthy. Row stays a real public result but
|
||
is **excluded from the model-record denominator** (`getModelAggregate`), like
|
||
void/unrecoverable.
|
||
- `analysis_flags` → valid for settlement AND the record; excluded only from
|
||
per-game/opponent analysis. **Never filtered from aggregates** (a test asserts
|
||
this, so a future edit can't silently collapse the two).
|
||
|
||
**Corrections to my own Phase 0 numbers, found while executing:**
|
||
- The 50 WNBA "poisoned moat rows" were **NOT poisoned** — captured 03:02 UTC
|
||
(23:02 ET 07-19) for game_date 07-19, so ESPN's "today" was the right day.
|
||
- The 166 doubleheader rows came from **loose team matching** and could not be
|
||
reconciled (a strict predicate found 121). I flagged only the **70** rows the
|
||
FEED itself marks (`G2…`/`(Game1)`) — and those are attributable, not
|
||
unattributed. The wider unattributed set needs a stricter pass; it is advisory
|
||
metadata that affects no number today.
|
||
|
||
**Final state:** 360 hit · 218 miss · 64 void · 55 pending · 25 quarantined
|
||
(697 public rows). 26 pending mlb 07-20 rows left untouched to self-heal on the
|
||
next cycle, as ordered.
|
||
|
||
## 🟢 GRADING DATE BINDING FIXED — wrong-opponent bug closed (2026-07-20 07:50 UTC)
|
||
`computeFeatures.lookupTodayGame` → **`lookupGameOnDate`**: it now sends
|
||
`?dates=YYYYMMDD` from the prop's BOUND game instead of taking ESPN's dateless
|
||
"today". Grading, settlement, retention and the ledger finally reference the
|
||
SAME game. **This was a MODEL-OUTPUT bug**: `opponentAbbr` feeds `opp_rank_stat`
|
||
(a ±1.0 factor) and `isHome` feeds `home_away` (+0.5).
|
||
|
||
**PROVEN against live ESPN, same instant:**
|
||
```
|
||
dateless "today" CLE→PIT NYY→LAD LAD→NYY (Jul 19 card)
|
||
bound to 2026-07-20 CLE→MIN NYY→PIT LAD→PHI (the real games)
|
||
bound to 2026-07-19 CLE→PIT NYY→LAD LAD→NYY (reproduces OLD exactly)
|
||
```
|
||
**Every opponent was wrong.** Measured at 07:5x UTC = **03:5x ET** — ESPN's
|
||
dateless scoreboard was STILL serving the previous day's card, so **the affected
|
||
window is wider than the 01:00/03:00 slots**.
|
||
|
||
- **Honest degradation:** no bound date → NO fallback lookup. Records
|
||
`no_bound_game_date`, leaves `opponentAbbr`/`isHome`/`gameId` null, and engine1
|
||
omits those factors rather than scoring a wrong matchup. Tests lock both ways.
|
||
- **Same class, also fixed:** Tank01 augmentation keyed its cache on TODAY's UTC
|
||
date; now uses the bound game date.
|
||
- **Audited:** no other dateless/"today" lookups remain in the feature path
|
||
(weather is current-conditions by venue; park/pace are static).
|
||
- Live cycles at 07:50/07:51 UTC wrote retention rows with `game_date 2026-07-20`
|
||
and `home_away` on 38/50 — games binding. MLB `opp_rank_stat` stays 0 for the
|
||
separate, known reason (ESPN's MLB team endpoint carries no defensive metric).
|
||
|
||
### 🔴 BLAST RADIUS (report-only — gates the heal)
|
||
**`ledger_entries` — grades by UTC hour (at-risk = late slots):**
|
||
|
||
| UTC hour | grades | settled results | risk |
|
||
|---|---|---|---|
|
||
| 00 | 32 | 29 | 🔴 at risk |
|
||
| 01 | 68 | 67 | 🔴 at risk |
|
||
| 03 | 152 | 103 | 🔴 at risk |
|
||
| 06 | 51 | 0 | 🔴 at risk (ESPN still on prior card at 02 ET) |
|
||
| 13/14/19/22 | 390 | 379 | ✅ same-ET-day, correct |
|
||
|
||
**At-risk: 303 grades (199 already settled with real results).** ~44 % of the
|
||
public record was graded in the affected window.
|
||
|
||
**`model_snapshots`:** 350 rows total — **50 carry `opp_rank_stat`** (the 03:02
|
||
UTC WNBA cycle) and are the poisoned moat rows; 190 carry `home_away`. Rows
|
||
captured 07:50+ are clean (post-fix).
|
||
|
||
**Recoverable vs quarantine (PLAN ONLY, nothing written):**
|
||
- **RECOVERABLE:** the correct opponent for a past date is still fetchable from
|
||
the date-pinned ESPN scoreboard, so a grade can be RE-DERIVED — but only where
|
||
the feature vector is reproducible. `model_snapshots` rows store the full
|
||
feature vector, so those are re-computable; **ledger rows do not** and would
|
||
need a full re-grade against historical inputs.
|
||
- **QUARANTINE:** any at-risk ledger grade that cannot be re-derived should be
|
||
marked computed-against-wrong-opponent and EXCLUDED from harness/validation —
|
||
never silently trusted. It stays in the public record as a real settled bet
|
||
(the result happened) but must not train or validate anything.
|
||
- **Scope note for the heal:** at-risk is defined by grading HOUR, not by proven
|
||
wrongness. Proving per-row wrongness requires re-deriving each row's opponent —
|
||
that IS the heal work.
|
||
|
||
## 🟢 game_date ROOT FIXED + 64 WRONG VOIDS REVERTED (2026-07-20 07:30 UTC)
|
||
- **ROOT (not DNP, as first diagnosed):** PropLine emits **no `commence_time`**
|
||
(grep-verified), so `ledgerService`'s `dateET(prop.game_time) ||
|
||
dateET(gradedTs)` always fell through to the **GRADE clock**. A 01:00/03:00
|
||
UTC snapshot is 21:00/23:00 ET the PREVIOUS day → tonight's props filed under
|
||
yesterday → settlement correctly found no game → Order 1's void logic
|
||
destroyed 64 real results.
|
||
- **FIX — `src/services/gameBinder.js`:** binds each prop to a scheduled game by
|
||
TEAMS across the ET window (grade date, +1, −1) and attaches the GAME'S own
|
||
time/date/id, in `snapshotService` **before** grading and the ledger write, so
|
||
ledger + retention + settlement all inherit one correct date.
|
||
**PROVEN against the live schedule at the failing hour:** a `2026-07-20T03:00Z`
|
||
grade (ET date 07-19, what the old code used) bound Steven Kwan to
|
||
**game_date 2026-07-20**, `mlb:2026-07-20:401816188`.
|
||
- **HARD CONTRACT:** unbindable → **nothing**. The grade-clock fallback is gone;
|
||
ledgerService SKIPS and counts such rows. Verified live: a prop with an
|
||
invented matchup returned `unresolved`, never a guessed date.
|
||
- **Doubleheaders** are flagged `ambiguous`, never attributed. Real case already
|
||
in data: `mlb:2026-07-11:MilwaukeeBrewers@PittsburghPirates(Game1)`.
|
||
- **Retention had the same bug** (rows dated by the snapshot clock) — fixed to
|
||
the bound game's ET date.
|
||
- **64 wrong voids REVERTED** (60 mlb + 4 wnba), scoped exactly to
|
||
`settlement_version=2 AND outcome='void' AND settlement_source='player_dnp'`.
|
||
- **LIVE PROOF of the safety fix:** the next induced settle pass returned
|
||
**`voided: 0, pending: 64`**. The old code voided all 64; the deployed code
|
||
refuses to claim DNP on an unconfirmed absence.
|
||
|
||
### 🔴 STOP-AND-REPORT — the date bug also reaches GRADING (not just settlement)
|
||
`computeFeatures.lookupTodayGame` calls the ESPN scoreboard with **no date
|
||
param** — it takes whatever ESPN calls "today". A prop graded at 23:00 ET can
|
||
therefore bind its `home_away` / `opponentAbbr` features to the WRONG game
|
||
(yesterday's finished one), which feeds `opp_rank_stat` and the home/away factor.
|
||
**Not fixed — flagged per the order's stop rule.** Blast radius:
|
||
- **Grading: AFFECTED** (feature correctness at late-ET slots — 01:00/03:00 UTC).
|
||
- **Settlement: FIXED** (this order).
|
||
- **Retention: FIXED** (this order).
|
||
- **Display/slate: NOT affected** — the board renders from the date-pinned
|
||
schedule and joins grades by player+stat, never by ledger `game_date`.
|
||
|
||
### 🔶 STILL OPEN — the reverted 64 cannot settle yet
|
||
Their `game_date` is still the OLD wrong value (Jul 17/18); the binder only
|
||
corrects rows written from now on. They will retry and age to `unrecoverable`
|
||
until their dates are corrected. **Correcting them is the first act of the heal
|
||
order** — healing against wrong dates would simply re-fail.
|
||
|
||
## 🟢 RETENTION CLOCK TICKS ON ITS OWN (2026-07-20 06:05 UTC)
|
||
Induced the REAL cron entrypoint rather than waiting for 14:00 UTC.
|
||
- **Cron capture needed NO wiring — verified in code.** `snapshotScheduler`'s
|
||
tick calls `runAll` = `snapshotService.runAllSnapshots` → `runSnapshot` per
|
||
sport → `gradeAndCacheSlate` → `onGraded` → retention. **The scheduled and
|
||
manual paths are the same function.** Nothing had been captured only because
|
||
no slot had fired since retention deployed (~02:55; slots 14/19/22/1/3).
|
||
- **PROVEN by induction:** `POST /api/internal/snapshot/all` (= `runAllSnapshots`,
|
||
the exact function the tick calls) wrote cycles at **06:03:23 mlb** and
|
||
**06:05:09 wnba**, 50 rows each, `cycle_hour_utc=6`, features 50/50.
|
||
*(The HTTP request gateway-timed-out at ~2 min; the job completed server-side —
|
||
verified in the DB, not from the response.)*
|
||
- **`archetype` FIXED — 50/50 on MLB** (was 0/100). Retention still COLLECTS at
|
||
grade time (only moment features exist) but now PERSISTS after enrichment via
|
||
`retentionService.mergeEnrichment`, which fills ONLY archetype/team/opponent
|
||
and never rewrites grade-time features or model output (test-locked).
|
||
- **🔶 `team`/`opponent` STILL NULL — upstream, not the merge.** The ENRICHED
|
||
grades themselves carry `team:null` (verified live: 0/25 with team, 25/25 with
|
||
archetype). `resolvePlayerStats` doesn't return `stats.team` for MLB, and
|
||
enrichment never sets `opponent` at all. Fixing that is an upstream change,
|
||
deliberately not made under this order.
|
||
- **WNBA archetype is null** — honest absence: WNBA stats don't resolve (Python
|
||
service offline), so there is nothing to classify from.
|
||
- **ZERO-WRITE ALARM ARMED** — `opsWatch.retentionZeroWriteAlarm` pages at
|
||
missed-snapshot severity when a slot GRADED props but retention wrote fewer
|
||
rows than the slate. `runSnapshot` now returns `retentionRows` so the tick can
|
||
evaluate it. Retention is best-effort by design, so a broken write is silent by
|
||
construction; this is the counterweight. Verified quiet on a healthy run
|
||
(gradeCount 25 → retentionRows 50).
|
||
- **INTRADAY DEFERRED (Phase 4).** `intradayRefreshService` calls
|
||
`analyzeViaEngine1` DIRECTLY — it never goes through `gradeAndCacheSlate`, so
|
||
there is no `onGraded` hook to reuse. Capturing it needs its own ctx +
|
||
snapshot_id + persist path and would ~5× row volume. Not small; not done.
|
||
- Outcome stamping still NOT implemented (depends on the settlement fix).
|
||
|
||
## 🟢 RETENTION LIVE — model_snapshots capturing (2026-07-20, migration 025 applied)
|
||
Phase 2 priority zero. **History now compounds from tonight.** Spec:
|
||
`specs/snapshot-retention.md`.
|
||
- **Why:** the ONLY model history was `ledger_entries` (640 rows / 6 game days)
|
||
and it stores NO model inputs. All ~30 other warehouse tables are empty. We
|
||
could score our own grades but could not replay a different model — the only
|
||
question a backtest answers.
|
||
- **`model_snapshots`** (migration 025, APPLIED): append-only, ONE ROW PER GRADED
|
||
PROP PER SIDE PER CYCLE. Unique on `(snapshot_id, player_key, stat, line,
|
||
side)` so a retried cycle can't duplicate. RLS on, service-role writes only.
|
||
- **VERIFIED CAPTURING (2026-07-20 03:02 UTC, 2 triggered cycles): 100 rows.**
|
||
MLB 50 (14 graded / **36 refused**), WNBA 50 (50 graded). `features`,
|
||
`grade_11`, `p_win`, `ev_pct` populated on 100% of graded rows; 1 model_version.
|
||
- **Refusals are captured — the ledger drops them entirely.** First-ever refusal
|
||
breakdown: `juiced_no_edge` 18, `rare_event_over_below_line` 13,
|
||
`insufficient_data` 5. **The MLB gate refused 36 of 50 sides (72%)** — whether
|
||
that costs edge is now measurable for the first time.
|
||
- Contract: retention is best-effort + injectable; `persist()` returns errors and
|
||
never throws. **A retention failure can never break a snapshot.**
|
||
- 🔶 **EV LOOKS OVERCONFIDENT — do not surface it as truth yet.** First captured
|
||
EVs include +62.1%, +61%, +56.9% — implausible for real markets. Cause is the
|
||
quantile estimator clamping `p_win` at 0.95 (`PROB_CEIL`) off ~10 games.
|
||
**Hero v2 already ranks by highest `ev_pct`, so it will select the MOST
|
||
overconfident read.** Calibration (C-cal) must run on the harness before EV
|
||
drives anything user-facing.
|
||
|
||
## 🔶 SETTLEMENT-CORRECTNESS AUDIT — two concrete findings (open, scheduled)
|
||
Logged from the Phase 2 history report; do NOT fix inside a retention build.
|
||
1. **ZERO pushes across 470 settled rows.** Implausible — hits/TB land exactly on
|
||
the number regularly. The exact-number push case is almost certainly
|
||
mis-scored (silently counted as hit or miss).
|
||
2. **~28 props/day never settle** (Jul 17 MLB 86 graded/57 settled; Jul 18
|
||
103/75). Cause undiagnosed.
|
||
Both are correctness bugs in the scoring that accuracy, ROI, CLV and every public
|
||
record depend on.
|
||
|
||
## 🔴 MODEL-VERSION CONTAMINATION in ledger_entries (permanent, mitigate by cutoff)
|
||
`ledger_entries` mixes pre- and post-2026-07-19-fix grades with **no version
|
||
marker**, and the eras cannot be separated retroactively. **Any backtest or
|
||
accuracy claim computed off existing ledger history must treat the fix boundary
|
||
as a hard cutoff.** `model_snapshots` stamps `model_version` + `code_sha` on every
|
||
row so this can never recur.
|
||
|
||
## 🟢 BACKUP — DURABLE ON-BOX, VERIFIED BY READ-BACK (2026-07-20, deployed)
|
||
The highest-severity open item is closed **on-box**. Off-box is a tracked follow-up.
|
||
- **Runs as CODE, not a cron somebody must install** (`src/backupScheduler.js`,
|
||
armed in `server.js`). Arming is opt-OUT: armed whenever `SUPABASE_DB_URL`
|
||
exists, `BACKUP_CRON=0` kills it, nightly `BACKUP_HOUR_UTC:BACKUP_MINUTE_UTC`
|
||
(default 03:10 UTC). The S62 design was opt-in and nobody ever opted in — that
|
||
is why the DB went unbacked for weeks. Deploy == installed now.
|
||
- **DURABLE:** `BACKUP_DIR=/app/backups` is a Coolify persistent volume, so dumps
|
||
survive redeploys. The container-ephemeral risk is closed.
|
||
- **VERIFIED END-TO-END (2026-07-20 02:31 UTC):** `POST /api/internal/backup/run`
|
||
→ exit 0 in 31s → `vyndr-20260720-023136.dump`, **821,549 bytes**, 698 archive
|
||
objects, ledger_entries present. `GET /api/internal/backup/verify` read the
|
||
rows back OUT of the archive with `pg_restore`: **643 rows in the dump vs 643
|
||
live — exact match**, DB 18 MB.
|
||
- The earlier figure of 604 was a moving target, not a discrepancy: the ledger
|
||
is append-only and 39 rows were written between the census and the dump
|
||
(2 manual regrades + the 01:00 UTC cron; newest lock 01:01:45 UTC).
|
||
604 + 39 = 643 reconciles exactly.
|
||
- **Permission bug found and fixed by the real run:** the first attempt failed
|
||
with `pg_dump: Permission denied` — the container runs as non-root `vyndr`
|
||
(uid 100/gid 101) and the mounted volume was root-owned. The Dockerfile now
|
||
creates + chowns `/app/backups`; the verify endpoint reports `uid`/`gid`/
|
||
`backup_dir_writable` so a mounted-but-unwritable volume is loud, not silent.
|
||
- **🔶 OFF-BOX IS DEFERRED, NOT DONE.** Storage Box SSH auth is not working yet,
|
||
so the rsync is gated behind `BACKUP_OFFBOX=1` (+ `BACKUP_REMOTE` +
|
||
`BACKUP_SSH_KEY`) and currently logs "DEFERRED". A failing push is a LOW
|
||
priority notice, never a failure — the durable on-box dump succeeded, and
|
||
crying wolf trains us to ignore backup alerts. **Remaining risk: a single
|
||
Hetzner-side loss takes the dumps with it.** To finish: install the generated
|
||
public key on the Storage Box, set `BACKUP_SSH_KEY` in Coolify (port 23 —
|
||
22 is SFTP-only), flip `BACKUP_OFFBOX=1`.
|
||
- **Not yet done:** a full restore into a scratch Postgres SERVER. The read-back
|
||
proves the archive CONTAINS the data; a server restore would additionally
|
||
prove schema/constraint restorability. It needs the dump off the container,
|
||
which is blocked on the same Storage Box auth.
|
||
|
||
## MODEL TRAIN — ARC 2 PREP: DEAD PROBABILITY LAYER REVIVED (2026-07-19, on main, DEPLOYED + fingerprinted)
|
||
Re-sequenced ahead of the gate work after the audit found the ranking signal did
|
||
not exist. Reports: `specs/audit-data/gate-simulation.md` (G-b/C-cal) +
|
||
`specs/audit-data/grade-collapse.md` (the big one).
|
||
- **p_win/ev_pct/kelly/model_odds/value were absent on 100% of live grades** —
|
||
`gameLogService` returns null for MLB by construction + the Python service is
|
||
offline, so `meta.gameLogs` was `[]` everywhere. EV, the Model Train's whole
|
||
ranking signal, had never been computed on a live prop; hero v2 always fell
|
||
through to its fallback; quarter-Kelly (sold on /pricing) never ran.
|
||
FIX: `featureCache.getStatRows`. **Fingerprinted live: p_win 0.523, ev_pct
|
||
-10.4, model_odds -109 — the value triplet is whole.**
|
||
- **Grade range restored ON MERIT** (rescaling permanently rejected):
|
||
`refreshTeamStats` wired into `runSnapshot` (it had ZERO callers), L20 made
|
||
symmetric (no negative branch existed), `game_count_in_7d` derived. D now emits
|
||
(WNBA 1/25, earned at p_win 0.365). **A-emission awaits the 01 UTC snapshot** —
|
||
`opp_rank_stat` needs that run. MARKETING HOLD on "A-RATED" copy until then.
|
||
- **Calibration guard:** consistency CV is NBA-tuned; `cv ~ 1/sqrt(mean)` made
|
||
every MLB stat `boom_bust`. Floored at `CONSISTENCY_MIN_MEAN=4`.
|
||
- Ledger truth (576 rows): only B/C ever emitted, 9 confidence values, ceiling 63;
|
||
past -400 was -7.7% ROI (the -400 floor was the whole win); flex -161..-250 is
|
||
our BEST band (+2.2%) so its EV gate ships DISABLED. Suite 276/3286, build 0.
|
||
|
||
## MODEL TRAIN — ARC 1 (ENGINE) SHIPPED (2026-07-19, on main `7a925f4`, pushed, NOT yet deploy-fingerprinted)
|
||
Spec written retroactively: `specs/model-train.md` (arc list + what's open). Goal of the
|
||
train: make "real opportunities at takeable prices" the ENGINE, not a display filter.
|
||
Arc 1 = steps 1–6, backend only. **No frontend consumes the new fields yet** — the
|
||
triplet/VALUE marker/refusal copy are Design's Session-2 surfaces.
|
||
- **De-vig** (`src/utils/devig.js`, NEW): two-way multiplicative de-vig → fair prob +
|
||
fair American price per side + `overround`. One side unpriced → fair is NULL, never
|
||
faked. Also `americanToImpliedProb` / `americanToDecimal` / `impliedProbToAmerican`.
|
||
- **EV** (`devig.evPct`): `ev_pct = modelProb × decimal − 1`, at the graded side's
|
||
ACTUAL price. This is the ranking signal now, replacing raw |model − consensus|.
|
||
- **Takeable / value gates** (`src/config/valueEngine.js`, NEW): `TAKEABLE_ODDS_CEILING`
|
||
**−160**, `TAKEABLE_ODDS_MAX` **+200**, `VALUE_EV_THRESHOLD` **2** (%) — all env-tunable.
|
||
`takeable` = price in band; `value` = takeable AND ev ≥ threshold. Strict null guards
|
||
(`Number(null)===0` would have made a missing price "takeable").
|
||
- **Hero v2** (`heroPropService`): the hero is the highest `ev_pct` among **takeable**
|
||
A/B reads — a huge model-vs-line gap on a −900 line is trivia, not an opportunity.
|
||
Falls back to most-recent-real graded read on an empty slate (unchanged).
|
||
- **Value triplet** on every graded read: `book_odds` · `fair_odds` · `model_odds`
|
||
(+ `fair_prob`, `overround`, `devig_method`). Wired inside `analyzeViaEngine1`'s
|
||
existing p_win/Kelly block; rides the `...result` spread through snapshot / hero /
|
||
scan. Documented in `BACKEND_HANDOFF.md`.
|
||
- **Test baseline: 276 suites / 3306 tests green** (was 274/3289), `next build` exit 0.
|
||
New suites: `devig.test.js`, `valueEngine.test.js` (+ `heroPropService.test.js` rewritten).
|
||
- **Knobs NOT in this arc** (asked about, confirmed absent from the codebase):
|
||
`EDGE_FLEX_WALL`, `HARD_JUICE_WALL`, `LADDER_ODDS_MAX`, `MIN_RUNG_PROBABILITY`.
|
||
The only price wall on disk is the flat `JUICE_ODDS_FLOOR` (−400) from the
|
||
rare-event work below — it is **not** edge-aware, and there is no −250 wall.
|
||
|
||
## RARE-EVENT UNDER SUPPRESSION (2026-07-19, on main)
|
||
Betting-logic audit: the CONSENSUS-vs-MODEL board flooded with "doubles u0.5 · MODEL 0.2 · +edge" fake reads (juiced rare-event unders). Report: the doubles projection is REAL per-player (MLB_LOG_FIELD doubles→doubles; values varied 0.03/0.16/0.2/0.22) — NOT a flat fallback; the issue is purely structural.
|
||
- **Config-driven** (`src/config/rareEventMarkets.js`): RARE_EVENT_STATS = doubles/triples/home_runs/stolen_bases, RARE_EVENT_LINE_MAX = 0.5.
|
||
- **Grade layer** (`analyzeViaEngine1`): rare-event UNDER at ≤0.5 → always refused (grade null + suppressed); rare-event OVER at ≤0.5 → refused unless projection > line (a 0.2-over-0.5 carries the SAME |edge| as the under, so it'd just take its board rank — refusing it is what actually clears the market). OVER with a genuine projection > line still grades.
|
||
- **Board layer** (`marketBreadth.collectBreadth`): drops null-model rows — a suppressed/ungraded prop can't rank a "MODEL —" placeholder onto the board.
|
||
- Suite 274/3289 green. Fingerprint: next MLB snapshot should have no rare-event u0.5 grades.
|
||
|
||
## BACKUP + FOUNDER CHECKOUT (2026-07-18, on main)
|
||
- **Task A — DB backup (SHIPPED; Kev cron+fingerprint on box)** `c2c43cd`: Dockerfile now has pg_dump/pg_restore/rsync; `backup-db.sh` validates every dump via `pg_restore --list` (must contain ledger_entries). Runs IN the API container (SUPABASE_DB_URL is there; WSL2 can't reach Supabase). Runbook = host cron `docker exec <api> sh /app/scripts/backup-db.sh`. **Mechanism fingerprint PASSED locally** (137 rows → dump → validate → restore → 137 rows). Kev: install the host cron + run the prod restore fingerprint.
|
||
- **Task B — founder checkout SEAT-GATED (SHIPPED)** `ccb9668`: `resolveCheckoutPrice` attaches the founder price while seats remain (< FOUNDER_SEATS_TOTAL, same `countFounderSeats()` truth as the meter), flips to standard at 100; meter shows SOLD OUT. Fixes "Claim a Founder Desk" charging $44.99 vs advertised $34.99. **payment_failed grace 48h→14d** (spans Stripe retries; revoke only on real cancel). Tested seat 0/99/100/null. Needs `STRIPE_PRICE_*_FOUNDER` set in prod for founder pricing to activate.
|
||
|
||
## SECURITY + PLUMBING FOLLOW-UP (2026-07-18, on main) — code shipped, migrations authored
|
||
Chrome ops session findings (Supabase advisor + Stripe). Code items shipped + fingerprinted; DB migrations authored for Kev to apply (no DB access from the box).
|
||
- **Item 0 (SHIPPED, fingerprinted)** `3b12c6c`: founder count = real active Stripe subs (`stripeService.countFounderSeats`), not DB tier fields. Phantom 1 (a comped profile) gone → prod now `available:false` (founder price IDs unset in prod env — set `STRIPE_PRICE_*_FOUNDER` to show the real 0).
|
||
- **Items 1,3,4,5 (AUTHORED — Kev applies in Supabase)** `78c1929`: `migrations/023_security_hardening.sql` (founder_pricing_seats → security_invoker; waitlist anon-INSERT-only; function search_path) + `migrations/024_anon_revoke_discoverability.sql` (revoke anon SELECT). Fingerprint = re-run Security Advisor.
|
||
- **Item 2 (SHIPPED — Kev sets env + cron)** `5f5c004`: `scripts/backup-db.sh` + `docs/BACKUP-RUNBOOK.md`. Needs `SUPABASE_DB_URL` (direct :5432 URI) in Coolify + the cron on the Hetzner box. Fingerprint = restore a dump + count ledger_entries.
|
||
- **Item 6 (SHIPPED — Kev activates portal in Stripe)** `889e862`: portal proxy + "Manage billing" button. Dunning verified: cancel-on-exhaustion → subscription.deleted → 48h grace → `gracePeriod.checkGracePeriod` revokes to free. Nuance: 48h grace on first payment_failed < Stripe's 2-week retry window (self-correcting).
|
||
- **Item 7 (SHIPPED)** `ae3cff9`: ESPN BET → theScore Bet (PENN) in books.js + ALLOWED_BOOKS; design HANDOFF flagged for the BookChip swap.
|
||
|
||
## TRUTH-EVERYWHERE PART 2 — COMPLETE (2026-07-18, on main, all 8 items fingerprinted)
|
||
Standing law: NO FABRICATED DATA ANYWHERE (UI, marketing, demos, examples). If it renders a number, it's from the DB or it doesn't render. Live data IS the better demo.
|
||
- **Item 1 — Verb law** (`66d52a9`): one verb, READ (never SCAN), swept every user-visible surface (FAB/nav/tier/ticker/upgrade). `tests/unit/verbLaw.test.js` enforces. Ticker "slate read" text flips at the next snapshot cron.
|
||
- **Item 2 — Founder counter** (`41fc2b9`): real active-paying-founder count from Stripe-synced user_profiles, cached; HIDES when unavailable. Live: **1/100** (was fabricated 47/100).
|
||
- **Item 3/4 — Reports**: full fabricated-data inventory + pricing claim audit (LIVE/PARTIAL). Alt-ladder/Kelly/correlation are LIVE in the engine; scratch-cascade + real-time-feed are PARTIAL.
|
||
- **Item 5 — Daily hero prop** (`9b9aab4`): live RULE = largest |model−consensus| gap, A/B only. Shows book line vs VYNDR model + real timestamp. Empty→most-recent-real; nothing→hidden. NO static Jokic fallback. `heroPropService` + `/api/hero-prop`. Live: Erica Wheeler 9.5 vs 14.2.
|
||
- **Item 6 — Desk showcase** (`cb3237c`): real alt-line ladder from the snapshot (edge guarded on small-line artifact); φ = real same-team correlation for two real legs or hidden; Kelly tile REMOVED (no odds → can't size honestly). `deskShowcaseService` + `/api/desk-showcase`. Live: Vandersloot ladder.
|
||
- **Item 7 — Accuracy → TRUTH** (`89a2977`): public 30D accuracy (`/api/accuracy` + `/api/ledger/accuracy`) now reads the CLEAN Postgres ledger aggregate (getModelAggregate/getAccuracyView, model_value>0), not the Redis outcome log. Number moved 67%→61%. **BEAT CLOSE suppressed** at source (`clvCaptureReliable()`) — see C4.
|
||
- **Item 8 — Blog** (`3b7a1f5`,`a8e383e`): wired the 5 REAL articles (content lives in `web/content/articles` — the runtime content root, NOT repo-root content/) with honest dates + markdown render + OG. Backdated 2026-03-22 orphan deleted + 301. Live: 5 articles serve.
|
||
|
||
### 🔴 C4 — CLV CAPTURE BROKEN (open, backend)
|
||
`captureClosing` re-records the LOCKED line as "closing" (closing_line == locked_line on the whole sample) → CLV is flat → BEAT CLOSE = a measured-wrong 0%. Suppressed on all public surfaces until fixed. Full write-up + fix plan: `specs/audit-data/clv-capture-broken.md`. Flip `CLV_CAPTURE_RELIABLE=1` after C4 verifies real closes.
|
||
|
||
## PHONE AUDIT PUNCH LIST — COMPLETE (2026-07-17, on main, deployed, fingerprinted)
|
||
Kev's 35-screenshot real-390px audit (signed-out + Desk signed-in). Worked top-down, fingerprinted the batch (compare "on current form" literal live in served chunk off `71e35e9`). Test baseline: 268 suites / 3239.
|
||
- **P0-1** blue-B/amber-C banished → token-derived grade badges (earlier batch).
|
||
- **P0-2** image pipeline (logos/headshots) `loading=eager` (earlier batch).
|
||
- **P0-3** player grouping/ladders across leaders/explore/ledger/breadth (`lib/playerGrouping.js`).
|
||
- **P0-4** mobile app bar collapsed to one line (ticker+heartbeat hidden <768, `MobileSyncClock`).
|
||
- **P1-5** 390px overflow containment (breadth rows stack, pitcher line wraps).
|
||
- **P1-6** (`ff0d3b1`) ONE freshness source: Slate "UPDATED" now reads the pipeline `refreshed_at` (the app-bar clock's field), not the client poll time — no more "0s ago" vs "SYNC 46:03" contradiction.
|
||
- **P1-7** (`d755b43`) flat edge board: an impossible `|edge|>40` is treated ABSENT at both data (can't rank a fake +140% over a real +8.4%) and display layers. Chips render now (P0-2). `EDGE_BOARD_SANE_MAX`.
|
||
- **P1-8** (`00537eb`) Compare verdict truth: no phantom "edge tonight" for July NBA(0) → "on current form"; cites what actually wins (scoring/boards/playmaking, not usage).
|
||
- **P2-9** (`77e8937`) leaderboard stat labels `web/src/lib/statAbbrev.js` (SB/ER/TB) + **FLAGGED** the backend grade degradation (below).
|
||
- **P2-10** (`71e35e9`) ledger read-card density (padding/margins tightened; archetype showDesc kept — it's inline, zero height cost).
|
||
|
||
### 🔴🔴 BACKEND FINDING — MLB GRADE PIPELINE DEGRADED (needs backend work, NOT frontend)
|
||
Diagnosed against live `/api/snapshot/mlb` (2026-07-17). Full write-up + do-not list: `specs/audit-data/mlb-grade-degradation.md`. Frontend guards (sane-edge cap, statAbbrev) are damage control, NOT the fix.
|
||
- **projection == 0 for 9/25** grades — the MLB feature path feeds 0 instead of refusing (S58 `insufficient_data`). This also breaks `computeEdge` → the broken edge_pct.
|
||
- **edge_pct on a broken scale** {20,60,100,140} — not a market %. Real prop edge is single-digit.
|
||
- **grade↔confidence mismatch on 25/25** — B shown at conf 55 (= the B- band per `grade_thresholds.json`); systematic one-sub-tier gap. The surfaced `confidence` is not the probability that derived the letter (likely the data-sufficiency penalty applied to display-only). This is the audit's "B at 45%" — NOT a frontend labeling or threshold bug; DO NOT re-letter/re-threshold on the frontend.
|
||
|
||
## DESIGN COMPLETION TRAIN — SUBSTANTIALLY SHIPPED (2026-07-16, on main, deployed) · ONE MASTER CHROME AUDIT PENDING
|
||
Goal: the live product matches Design's mockup at EVERY width. **HONESTY LAW: I cannot see rendered output (WSL2↔Chrome unreachable) — every visual change is "built to Design's spec, VISUALLY UNVERIFIED." Kev runs ONE master Chrome audit at the END (390px + desktop) against the running unverified list below; fixes propagate through tokens (surfaces inherit, don't fork).** Design package (3 revs) authoritative at `specs/design-reference/` (HANDOFF.md = entry point; `vyndr-system.html` desktop, `vyndr-mobile.html` 19 screens, `vyndr-landing.html`, `assets/glyphs/` 83 SVGs+MANIFEST).
|
||
SHIPPED + deployed:
|
||
- **Palette alignment** (`e0a26a8`,`c3d469e`,`ebaab00`): tokens → Design's EXACT ramp. **Grade colors corrected: B blue→WHITE #F0F0F0, C amber→GREY #B8BCC8, D #FF4757** (the old blue/amber violated DESIGN-SPEC's own "B neutral-bright, C muted"). Surfaces #0E0E14/#14141E, text #F0F0F0/#B8BCC8/#707080, borders #1E1E2A/#2A2A38. vyndrTokens + tests aligned. ShareCard/OG billboards grade-B white.
|
||
- **Glyph system** (`77d8fd6`,`771d8b0`): 44 archetypes render Design's REAL marks (83-mark set: 74 display + 9 classifier-legacy), front `lib/archetypes.js` + back `archetypeService.js` color-synced. ZERO generic placeholders left. Combat keeps its char-glyph namespace (separate test).
|
||
- **Entity layer VERIFIED already-complete** (DS0/Wave 2A/2B — the audit was stale): TeamLogo (real ESPN CDN, monogram fallback) on game cards/team hub; PlayerAvatar headshots MLB/NBA/WNBA + NFL/NHL resolver ready + soccer-monogram (blocked) + combat-no-photos; books via bookInfo/BookChip everywhere (fixed BookComparison holdout `1558710`); line-shopping surfaces (detectBestBook/MarketBreadth) self-deepen. No `$1M` copy.
|
||
- **Rev 3 matchup chips** (`f24f941`,`baf977f`,`bf8bf16`): `TeamChip` primitive (real TeamLogo at 10-12px + abbr, inherits row opacity) on StatStrip board rows, grade reveal, pitcher-vs. Game-card headers already had TeamLogo.
|
||
- **Mobile: M1a header collapse** (`a10227e`) · **M1b app-bar clock** (`7c7ab24`, Hybrid: wall clock rests, STALE reacts off refreshed_at) · **grade-reveal mobile** (`baf977f`, 74px hero) · **flat EDGE BOARD** (`8a24ac9` — Design's core mobile screen: NEW flat edge-ranked IA, `slateAdapter.flattenToEdgeBoard`+`MobileEdgeBoard`, shown <768px, cards ≥768px; threaded edge_pct through the pipeline) · **M4 structural locks** (`6bc9093`).
|
||
- **Key finding:** the 19 mobile screens were LARGELY already built (pages built to the SAME DS0-DS5/§7 contracts Design's mobile follows) → they already match structurally + stack on mobile. The genuinely-new work was the flat board (done) + narrow chip/hero polish.
|
||
REMAINING (all VISUAL, for the master audit): per-screen mobile polish of the ~17 already-built screens (blind, mostly Design-aligned) · M2 billboard crop-ratio verification (OG routes exist, colors fixed). Test baseline: 266 suites / 3214.
|
||
|
||
### 🔴 MASTER-AUDIT UNVERIFIED LIST (every visual change this train — Kev's one end-of-session audit)
|
||
1. **Palette** (biggest): grade B-white/C-grey/red-#FF4757 + full surface/text/border ramp — every surface inherits.
|
||
2. **Archetype marks + colors** — 44 archetypes render Design's real 83-mark set.
|
||
3. **Flat EDGE BOARD** at 390px — the core mobile screen (rows/opacity-ramp/chips/hero-edge%).
|
||
4. **Grade-reveal mobile** (74px hero + team chip) · **app-bar clock** (Hybrid) · **M1a header collapse**.
|
||
5. **TeamChip** on board rows / grade reveal / pitcher-vs (real logo at 10-12px).
|
||
6. **OG billboards** (ShareCard/player/u) grade-B white + text #F0F0F0.
|
||
|
||
## SHIPPED — HEARTBEAT + CADENCE + QUOTA-GUARD (2026-07-15, on main, deployed)
|
||
Three ships, each green + deployed. Standing authority is now PERMANENT full autonomy (three-gate floor: tests green, web build exit 0, post-deploy fingerprint of NEW code).
|
||
- **`7712f0a` Heartbeat honesty (fingerprint-verified live):** the "frozen snapshot / STALE 8h" was a MEASUREMENT bug, not a dead pipeline. `updated_at` = grade-LOCK time (advances 5×/day by design; grades never change in-game); the SYNC badge measured the 20-min `expected_interval_s` against it → structurally STALE. `refreshed_at` (the real 20-min intraday heartbeat) was written to Redis but NEVER serialized. FIX: full snapshot seeds `refreshed_at`, `/api/snapshot/summary` + GET `/:sport` expose it, `LiveLayer` badge measures from it. Fingerprint PROVED the pipeline was alive all along (`refreshed_at` advancing 20:20→20:40→21:00→22:00; the 22:00 full slot fired → scheduler healthy, re-armed fine after the 19:37 restart).
|
||
- **`4cd933d` + `2d413cf` Per-sport cadence (Job 1):** `src/config/sportCadence.js` (config-over-constants). REALITY-CORRECTED: mlb/nba/**wnba** all keep the full grid `14/19/22/1/3` (PropLine 9k/day = cheap; WNBA games span afternoon→late-evening, so 14 UTC/10am ET catches afternoon-game props — the earlier "drop it as waste" was wrong, corrected by observing 2 finished afternoon WNBA games on Jul 15). **Soccer alone is trimmed** to `14/19` + NO intraday — the genuine win, protecting the scarce odds-api key. Scheduler fires at HOURS_UTC and grades only `sportsForHour(h)`; only `intradaySports()` refresh.
|
||
- **`2d413cf` Quota guard (the real finding):** odds-api quota was EXHAUSTED 0/500 mid-cycle, unpaged. Root cause: `futuresService` (the only regular burner) used raw axios, BYPASSING the gateway → never hit `recordCall` (the ONLY place the WARN/BLOCK pager fires) → the counter reached 100% via silent header-sync. FIX: futures now routes through `gateway.fetch('odds-api', …)` (counted + blocked); a RESERVE floor (`ODDS_API_RESERVE`=50) refuses DISCRETIONARY calls (futures/soccer) while `remaining ≤ reserve` so they can NEVER starve MLB's essential prop-backup; `quotaTracker.syncFromHeaders` (authoritative) now ALSO fires the once-per-period alert; `POST /api/internal/quota/test-alert` test-fires the pager end-to-end. **Fingerprint caveat:** backend-internal, no unauthenticated public surface (`/api/internal/*` 401s for any path — auth runs before route-match); verify via authenticated test-alert (200=new/404=old) with the internal key.
|
||
- **odds-api 0/500 → BOTH soccer props AND futures are blocked on QUOTA, not code.** Kev's ruling: HOLD THE LINE (don't pay; zero-out-of-pocket). Soccer stays honest-empty; cadence is ready for the reset. Env: keep `API_FOOTBALL_KEY` (code reads it), delete inert `API_FOOTBALL_KEY2` (Kev handling). See memory [[odds-api-quota-soccer-blocked]]. Test baseline: 264 suites / 3195.
|
||
|
||
## SHIPPED — TRUTH-EVERYWHERE + OFFSEASON HUB TRAIN (on main, deployed)
|
||
Governed by `specs/truth-everywhere-train.md`. **The reframe that drove it: only MLB GRADED in prod (NBA/WNBA refused their whole slate — offline Python feature source, no fallback; soccer errored pre-grade). So grading, not settlement, was the first domino.** Standing deploy authority granted (green + fingerprint → merge + deploy, no ask). Waves DEPLOYED:
|
||
- **Wave 0 — NBA/WNBA GRADE (free ESPN gamelogs):** `espnStatsAdapter.getPlayerGameLog` feeds `featureCache` when the Python source is null → l5/l20 → props grade. **Verified live** (A'ja Wilson→B, Ionescu→C off real ESPN form). Endpoint `site.web.api.espn.com/apis/common/v3/sports/basketball/{nba|wnba}/athletes/{id}/gamelog`.
|
||
- **Resolver-hardening:** roster-index primary resolver (206 WNBA / 544 NBA) + `?season=` gamelog retry — the real fix for dual-league (WNBA+NCAA) players who got an empty gamelog (Collier/Brionna Jones now resolve). Broad coverage.
|
||
- **Wave 1 — NBA/WNBA SETTLE:** `defaultGetPlayerStats` routes nba/wnba to the ESPN gamelog in BOTH `outcomeService` + `ledgerService`; separate `NBA_BOX_KEY`/`NBA_COMBO` settle maps; final-honesty guard (never settle in-progress); `opsWatch` finals-gated zero-settle alarm + per-sport boot announce. accuracy:{sport}+by_tier light up automatically. **PROOF is next-day** (grades lock tonight → settle tomorrow → `accuracy:wnba`).
|
||
- **Wave 2 — NEVER-DARK HUB:** ESPN `/news` wire (`newsService` + `/api/news/:sport`) — **verified live (real headlines)**; `NewsWire`+`FuturesBoard` on `/explore` (offseason-aware, self-hide on empty, futures "TRACKED · NOT GRADED"). **KNOWN GAP: `/api/futures/:sport` returns `markets:[]` for ALL sports in prod** — code path looks correct (odds-api `/sports/{FUTURES_KEYS}/odds?markets=outrights`); needs a live odds-api probe (ODDS_API_KEY) to finish. Self-hides honestly meanwhile. `FUTURES_ENABLED=0` kill-switch; 12h TTL; 1 credit/refresh.
|
||
- **Also rode along:** `espnAthleteIndex` cross-sport headshot capture (headshot-coverage) — NBA/WNBA/NFL/NHL real headshots when in-season.
|
||
DEFERRED (locked decisions): **soccer** — grade+settle path, needs `API_FOOTBALL_KEY` (VALIDATED, Free 100/day, active to 2027 — set it in Coolify env `API_FOOTBALL_KEY`; code dormant until wired) + ESPN-soccer-free-first attempt; **MMA moneyline settle** (combat not yet in the grade/lock loop); **NFL/NHL** (dormant, in-season later); **futures debug** (odds-api probe). Test baseline now **263 suites / 3183**.
|
||
|
||
## SHIPPED — WIRING & DATA TRAIN (on main `3b1aa9f`, deployed 2026-07-13)
|
||
Governed by `specs/wiring-data-train.md` + `specs/combat-intelligence.md` + the global visual reference `specs/design-reference/vyndr-system.html` (build toward it; live wordmark kept). **All 6 waves shipped + deployed** (fingerprint-verified: /parlay, /fight, /u/vyndr live; house record 88-38/70% real). Real assets verified live (MLB/ESPN headshot CDNs, ESPN-MMA feed, Baseball Savant CSV).
|
||
- **Wave 1 — trust bugs:** billing renewal honest render (`billingDisplay.classifyRenewal` — no far-future placeholder); James Wood nameKey-collision fixed (`mlbStatsAdapter` teamHint disambiguation + streaks join-invariant); DeskShowcase "$1M terminal" → deadpan copy.
|
||
- **Wave 2 — sport-agnostic entity layer:** real player headshots threaded from ingestion (MLB MLBAM + NBA/WNBA ESPN athlete ids that were fetched-and-discarded) across slate/scan/hotlist/search/grade card; soccer = honest monogram (no free id); 8 self-authored SVG book wordmarks (`web/public/books/*.svg`, swappable for official art) + all 10 book keys resolve; team-logo abbr aliases. Storage: id on the `enriched` grade at `snapshotService` (zero new I/O).
|
||
- **Wave 3 — record by grade tier (Addition 2):** ONE shared `TierRecord` (`lib/tierRecord.js` + component) on dashboard + /u + ledger; per-tier W-L always, hit-% only at n≥20 per tier (gate stays in `getModelAggregate`).
|
||
- **Wave 4 — missing surfaces:** Outlook mode (grid never blank → yesterday receipts / tomorrow schedule); Market-Breadth median-consensus-vs-model strip (self-hides <2 books); Parlay Lab `/parlay` (slate-independent leg source); live Grade-Shift timeline (`GradeShift`).
|
||
- **Wave 5 — /u house-mode + arsenal:** house handle `vyndr` surfaces the real `user_id=NULL` public model record + per-tier calibration + 1080×1350 portrait/OG (user-handle privacy 404s stay byte-identical); Baseball Savant pitcher-arsenal (`savantAdapter`, free CSV, verified) → `PitcherArsenal` card, self-hides on absent.
|
||
- **Wave 6 — combat v1 (MMA):** ESPN-MMA fight cards + tale-of-the-tape + style-blend archetypes (sport-scoped `COMBAT_ARCHETYPES` — FINISHER color collided w/ soccer + tripped the ΔE gate, so kept separate) + odds-api ML/round-totals + style-edge verdict. NOT in the snapshot/settle loop; method/round/props + fighter photos + matchup-GRADE engine + ufcstats scraping all DEFERRED, flagged data-limited in-UI. New routes: `/fight/[id]`, `/parlay`, `/u/[handle]/portrait`.
|
||
POST-MERGE TODO: NBA/WNBA headshot coverage + combat depth need prod runtime verification; soccer headshots blocked on `API_FOOTBALL_KEY`; combat settlement + matchup-grade engine are the next combat sub-wave.
|
||
|
||
## SHIPPED — DESIGN TRAIN v2 (merged to main `e9c0a59`, deployed & fingerprint-verified live)
|
||
Governed by `specs/DESIGN-SPEC.md` v2 (the raised standard: entities render as
|
||
themselves, color contract, sub-200ms, screenshot-first billboards). **All six
|
||
sessions DS0–DS5 shipped: merged to main, pushed to gitea, Coolify auto-deploy
|
||
LANDED — verified live via `/pricing` DeskShowcase ("A $1M terminal") + `/u`
|
||
route serving the unified EmptyState. 238 suites / 2885 tests green,
|
||
`next build` exit 0.** Post-ship TODO: re-run the Chrome design audit against
|
||
prod v2 to confirm all 22 audit findings closed (env here can't screenshot —
|
||
WSL2↔Windows-Chrome localhost unreachable).
|
||
- **DS0 Entity Layer** (`24af247`): `teamMeta.js` (real logos+colors for 30 MLB /
|
||
30 NBA / 13 WNBA / 48 WC nations via ESPN CDNs), `TeamLogo` / `PlayerAvatar`
|
||
(team-colored monogram fallback, no gray silhouette) / `BookWordmark`, swapped
|
||
into GameCard/StatStrip/StreaksPanel/TeamHub.
|
||
- **DS1 Speed + Trust Bugs** (`1c681df`): React #418 hydration fix (mounted-flag
|
||
guard), layout-matched `Skeleton`s replace text-wall loaders, scan→ledger
|
||
persistence via `session.access_token`. **DS1 follow-up** (`cf91c04`): closed
|
||
the sb-token trust-bug CLASS — `lib/authToken.js` `currentAccessToken()` reads
|
||
the REAL Supabase session (not the OAuth-only `sb-token` key) across
|
||
profile/slip/dashboard/settings/tracker.
|
||
- **DS3 Color Contract** (`49a3323`): `lib/colorContract.js` — green = one meaning
|
||
(edge/A-tier/CTA), edge/CLV by sign, glow A/A+ only, archetype greens deduped.
|
||
Enforced by `colorContract.test.js` (fails on violation).
|
||
- **DS4 Billboards** (`45bafbc`): STREAKS row (length as mono hero), grade reveal
|
||
(sign-colored edge), CLV reframe (flat says so), `/u` public profile + OG.
|
||
- **DS2 Dashboard Slate Rebuild** (`fe294a5`): never-empty hero (falls back to
|
||
yesterday's A-tier settled receipts), one bold hero per card, pending-filler
|
||
collapse, ranked tonight-grades. `slateAdapter.js` engine + 21 tests.
|
||
- **DS5 Pricing + Motion + States** (`a18a3f3`): Desk $44.99 as the hero tier
|
||
(single green CTA, `DeskShowcase`), ticker → punctuated stillness (≥4.2s hold,
|
||
one idle motion), unified `EmptyState` (404 grammar) at TeamHub/game/ledger,
|
||
archetype glyphs propagated to streaks/ledger.
|
||
NOTE: WSL2↔Windows-Chrome localhost is unreachable in this env, so visual proof
|
||
was SSR-render + source-assertion tests + live prod fingerprint curls, not
|
||
pixel screenshots. The `design/v2-train` + per-DS worktree branches can be
|
||
pruned. Stale `origin` GitHub remote still carries the long-dead leaked PAT —
|
||
deploy runs off `gitea`, so it's inert, but rotate it when convenient.
|
||
|
||
|
||
|
||
## WHERE THINGS STAND
|
||
|
||
**Tests: 2757 / 229 suites, all green. Web build exit 0.** Baseline at the A1
|
||
board's start was 2398. Governing docs: `specs/VYNDR-NORTH-STAR.md` (v1.1) +
|
||
`specs/VOICE.md` (v1.1) + `specs/ROW-GRAMMAR.md` — read them before building
|
||
anything.
|
||
|
||
**The record is live and writing itself.** Day one (Jul 11): 24→25 MLB rows,
|
||
real locks (line/odds/book), closing capture every snapshot + every 20-min
|
||
intraday refresh, settlement each morning slot with signed CLV. No percentage
|
||
renders anywhere under n≥20 — enforced in `ledgerService.getModelAggregate`,
|
||
the single gate. WNBA/soccer rows pend honestly until Phase 4.5.
|
||
|
||
## THE THREE A1 TRAINS (all merged to main, all pushed)
|
||
|
||
| Train | Range | Carried |
|
||
|---|---|---|
|
||
| #1 | `4d2b27d → 219167e` | S0 verify · S1 promise audit (`PROMISE-AUDIT.md`: alt ladder + quarter-Kelly BUILT, analyst unlimited, "40+" claims made honest) · S2 compliance (/responsible-gambling rebuilt, /terms /privacy drafts, /methodology, 5 seed articles in `content/articles/`) · S3 affiliate (BOOK IT deep links organic-until-config-flip, best-price dots, ?ref= partner attribution) · S4 media engine (`mediaEngine` + executable VOICE lint, `/desk` arsenal + DATA BRIEF, Ghost DRAFTS-only publisher) · S5 viability (statsapi lineups → CONFIRMED/NOT-IN kills grades visibly; ESPN injury wire; Yesterday/Today/Tomorrow nav + Settle panel) · S7 newsletter (Listmonk capture + THE VYNDR REPORT assembly, env-gated) · S8 ops (settle alarm, 3-slot failure pager, quota/disk/mem, 9AM pulse — test alert delivered to the real ntfy channel) · S9 slip reader (tesseract.js OCR + per-book parsers, /slip page; synthetic-image acceptance passed) · S10 public profiles (/u/{handle}, private-by-default, no-existence-leak) · migrations 021/022 |
|
||
| #2 | `219167e → 14dc9cf` | S6 display (ROW-GRAMMAR.md + slot-order fixes, line sparklines from intraday history, last-10 ●●○ dots, CLV distribution on MODEL tab, ⌘K SearchModal + mobile search, landing LCP fixes) |
|
||
| #3 | `14dc9cf → ec5ee3d` | S11 live tracking (statsapi/ESPN live box lines → ON PACE / NEEDS N / HIT ✓ in the outcome slot, live games float up, "TRACKING — READ LOCKED PRE-GAME"; real-live acceptance: Harper 3/1.5 TB ▼8th → HIT ✓). Grades never change in-game. |
|
||
|
||
**Deploy status:** ALL LIVE as of `17fb981` (Jul 12 17:43 UTC). Fingerprints
|
||
confirmed in prod: `/api/live/mlb` serves JSON with `hasLive:true` (S11 live
|
||
tracking working — a real game was tracked at confirm time), `/api/health`
|
||
healthy, `/api/snapshot/summary` → `expected_interval_s:1200` (env var now set).
|
||
|
||
**P0 postmortem (`17fb981`):** the API silently served a 14h-old image
|
||
(`4d2b27d`) from Jul 11 ~20:00 to Jul 12 17:43. Root cause: the Dockerfile
|
||
never copied `content/`, and `mediaEngine.js` read `content/stark-lines.json`
|
||
with an unguarded module-load `readFileSync` → ENOENT at require time crashed
|
||
`app.js`'s eager route chain (app → routes/desk → deskService → mediaEngine)
|
||
→ boot failed → Coolify healthcheck rolled back. Every merge from S6 onward
|
||
was merged-but-not-running for 14h. Fixed: Dockerfile COPYs content/;
|
||
mediaEngine treats the file as optional garnish (loadStark try/catch → {});
|
||
`src/preflight.js` now prints `[preflight] OK|DEGRADED` at boot so a missing
|
||
file/env is legible instead of masked by a rollback. LESSON: any module-load
|
||
file read is a boot-crash risk — keep them lazy + try/catch, and the
|
||
Dockerfile must copy every dir the runtime reads.
|
||
|
||
## LIVE INFRA MAP (verified in prod)
|
||
|
||
- **VYNDR Web** — Next.js :3000, vyndr.app. **VYNDR API** — Express :3001,
|
||
api.vyndr.app (`/api/internal/*` behind `VYNDR_INTERNAL_KEY`).
|
||
- **Scheduler (in-process, API app, `SNAPSHOT_CRON=1`)** — snapshot slots
|
||
14,19,22,1,3 UTC; settle pass (outcomes + ledger, idempotent) runs FIRST
|
||
each slot; 20-min intraday odds refresh during slate hours
|
||
(`INTRADAY_REFRESH=0` kills); desk-ready ping after the day's first slot;
|
||
daily pulse 13:00 UTC; missed-cron + 3-slot-failure + quota pagers → ntfy
|
||
`vyndr-pipeline-kev2026`. Boot logs: `[snapshotScheduler] armed` +
|
||
`[settlement] armed`.
|
||
- **Data**: Redis (mastermind-cache) for all pipeline caches (SNAP_TTL 24h —
|
||
do NOT shorten; the 11h overnight gap killed morning settles once).
|
||
**Supabase = the ledger** (`ledger_entries` + `public_profiles` +
|
||
`user_profiles.partner_ref`; RLS everywhere, service-role-only writes;
|
||
migrations 001–022 applied, files in `supabase/migrations/`).
|
||
- **Deploys**: gitea `git.builtbykev.com/builtbykev/vyndr` → Coolify webhook
|
||
(HMAC secret fixed Jul 11; auto-deploy proven). GitHub `origin` remote is
|
||
dead-PAT hygiene debt only.
|
||
- **Free feeds** (zero out-of-pocket law): PropLine ×3 keys (9k/day capacity;
|
||
intraday uses ≤144/day), odds-api backup (quota-gated), statsapi.mlb.com,
|
||
ESPN site API. tesseract.js OCR self-hosted (traineddata caches on first
|
||
call — needs outbound network once).
|
||
|
||
## ENV VARS STILL PENDING ON KEV'S SIDE (Coolify, API app)
|
||
|
||
1. `SNAPSHOT_EXPECTED_INTERVAL=1200` — **confirmed absent** (fresh restart
|
||
still served the 18000 default). SYNC badge tracks intraday cadence once set.
|
||
2. `DESK_OWNERS=kevdevelops@gmail.com` — /desk 403s everyone until set.
|
||
3. `LISTMONK_URL` / `LISTMONK_USER` / `LISTMONK_TOKEN` / `LISTMONK_LIST_ID`
|
||
(all four together) — newsletter no-ops politely without them; box-side
|
||
setup in `docs/NEWSLETTER.md`.
|
||
4. Optional: `GHOST_URL` + `GHOST_ADMIN_API_KEY` (drafts), `PULSE_HOUR_UTC`.
|
||
|
||
**Entity placeholders** (S2, blocked on Kev's values): `[ENTITY NAME]`,
|
||
`[STATE OF FORMATION]`, `[ARBITRATION VENUE]`, `[CONTACT EMAIL]` in
|
||
`web/src/app/terms/page.tsx` (6×) + `web/src/app/privacy/page.tsx` (5×).
|
||
A test keeps unverified entity names out until replaced.
|
||
|
||
## BOOK COMPARISON — data layer shipped (per the build order)
|
||
|
||
**What shipped (backend + tooling; UI routing is the gated next step):**
|
||
- **Phase 1 — fenced, snapshot-locked per-book store.** `src/services/
|
||
bookPriceStore.js` captures per-book prices from the multi-book `props` array
|
||
BEFORE `dedupeProps` runs, keyed by `nameKey|stat_type`, written to
|
||
`bookprices:{sport}` at SNAP_TTL (24h) inside `snapshotService.runSnapshot`.
|
||
STRUCTURAL FENCE: it only reads `props` and writes its own key; nothing on the
|
||
grade path reads it. Proven byte-identical — the graded slate is unchanged with
|
||
or without the capture (`tests/unit/bookPriceStore.test.js`, incl. a grep test
|
||
asserting no grade-path file references `bookprices`).
|
||
- **Phase 2 — the crown is threshold-gated.** `scripts/measure-book-spread.js`
|
||
reports best-vs-worst same-line price spread (American cents AND implied-prob
|
||
points), book-count histogram, %-zero, pinnacle presence — per sport, never
|
||
pooled. PRE-REGISTERED: crown ships only if median ≥8¢ OR ≥2 implied-prob pts.
|
||
Runs post-deploy against real captured data (Redis degraded locally → exits 0).
|
||
- **Phase 3 (backend) — honest-absent + gated crown.** `bookComparisonService.
|
||
compareProp` now renders single-book / flat-market props with NO crown; the
|
||
crown fires only among ≥2 books at the SAME line with DIFFERING prices, and only
|
||
when `BOOK_CROWN_ENABLED=1` (default OFF until Phase 2 clears). `/api/books/
|
||
:sport` (a crown claim) returns [] while gated off; `/api/books/:sport/:player/
|
||
:stat` returns the honest grid always. Route repointed to read `bookprices:*`
|
||
(fallback: odds cache), matched by `nameKey`. Response `source` field is the
|
||
deploy fingerprint (`bookprices` = new store serving).
|
||
|
||
**What it revealed:** the premise ("BookComparison.tsx is a dead component") was
|
||
incomplete — the ENTIRE feature (service + `/api/books` routes + Next proxy) was
|
||
built and mounted, but NON-FUNCTIONAL end-to-end: the route fed FLAT odds-cache
|
||
rows to a comparator that expects GROUPED `{lines|books}`, so it always returned
|
||
empty. Per-book prices were retained only transiently (odds cache, ~1h, raw
|
||
names, grade-path INPUT); every grade-path persistence point collapses to one
|
||
book. Fixed by the snapshot-locked store + route repoint.
|
||
|
||
**Gated next step (NOT this order-turn):** route `BookComparison.tsx` onto the
|
||
prop card fed by the store, and design the crown treatment — AFTER Phase 2's
|
||
measurement returns real spread numbers on prod (you can't honestly design a
|
||
crown before you know it ships). `BOOK_CROWN_ENABLED` stays 0 until then.
|
||
HELD unchanged: dedupeProps, snapshot dedup, the selector, the grade, champion,
|
||
challengers, ranking, edge_pct/ev_pct. No push-to-book, no movement strip, no
|
||
pinnacle edge test (data persisted only).
|
||
|
||
## HONEST OPEN ITEMS
|
||
|
||
- **Phase 4.5 — WNBA settlement via ESPN box scores. DUE ~Jul 24** (hard
|
||
date). Until then WNBA/soccer/NBA ledger rows pend forever; accuracy = MLB only.
|
||
- **The record needs time**: no percentage anywhere before n≥20 settles, by
|
||
design. First percentages expected ~day 3–4 of MLB volume.
|
||
- **Slip reader**: acceptance was a clean synthetic image; needs one real
|
||
phone screenshot (DK) to validate OCR robustness.
|
||
- **Landing LCP**: S6's fixes are static-analysis-provable but unmeasured —
|
||
run PageSpeed mobile vs the 3.8s baseline; next suspect is the
|
||
LiveHeroProp post-fetch swap.
|
||
- **Partner report endpoint** returns real numbers only for signups AFTER
|
||
migration 021 (metadata path live; OAuth signups carry no ref — known gap
|
||
in `docs/PARTNERS.md`).
|
||
- **NBA Oct / NFL Sep readiness** + soccer end-to-end: `specs/vyndr-roadmap.md`.
|
||
- **Killed permanently** (founder ruling): AI chat assistant, light mode.
|
||
|
||
## OPERATING LAWS (never relax)
|
||
|
||
Zero out-of-pocket · absent beats wrong, refusal beats hollow · n≥20 before
|
||
any percentage · nothing auto-posts anywhere, ever · grades never change
|
||
in-game (revisions are public via `revised_from_grade`) · one meaning per
|
||
color, data in mono · no exclamation points, anywhere · the pipeline is the
|
||
only source of numbers, including in marketing (`mediaEngine` lint enforces).
|
||
|
||
— STATE as of ec5ee3d · the record is writing —
|