diff --git a/outputs/VYNDR-COMPLETION-MATRIX.md b/outputs/VYNDR-COMPLETION-MATRIX.md index 579980e..0794101 100644 --- a/outputs/VYNDR-COMPLETION-MATRIX.md +++ b/outputs/VYNDR-COMPLETION-MATRIX.md @@ -426,3 +426,33 @@ are absent entirely. Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting Bayesian to Node) → wire Layer 2 → reconnect Layer 3 forward → Layer 1 → per-sport efficiency → challenger promotion → sport-as-module → NFL/CFB. MLB is specced as the reference module. + +--- + +# FULL-OUTPUT GRADE MAPPING + COLLAPSE COST — 2026-07-30 (report-only) → `specs/full-output-grade-mapping.md` + +Track-B 1 of 3. The three-layer engine is **BUILT but NOT WIRED and NOT DEPLOYED** (re-verified), so no +posterior/CI is produced today. Measured instead against the collapse that actually exists — three of +them: estimator components dropped; **`p_win` excluded from the grade entirely** (the severe one — the +letter is a factor index with zero probability input); and `grade_thresholds.json` read backwards to +manufacture `confidence`. Market-efficiency scaling is never computed at all. + +**THE MEASUREMENT (354 settled rows, locked pre-game p_win — forward, not lookahead):** + +| basis | ALL (n=354) | MLB (n=224) | WNBA (n=130) | +|---|---|---|---| +| champion letter → outcome r | **0.0050** (p≈0.93, null) | 0.0686 (n.s.) | −0.0986 | +| probability letter → outcome r | **0.1313** (p≈0.013) | **0.2356** (p≈0.0004) | **−0.1258** | + +**The served letter is INVERTED between its only two populated tiers — B 52.4% (n=168) vs C 56.9% +(n=174).** Probability letters spread 30.8%→70.0%, use 10-11 of 11 letters (champion uses 3-4), and +split ROI −1.42% (A-family n=78) vs −26.62% (C-/D/F n=51). + +**Verdict: costly on MLB, and un-collapsing does NOT help WNBA** (both correlations inverse there) — +so the full-output challenger must be **MLB-FIRST**. Five falsifiable mapping rules are specced (R2 +"uncertainty grades down" stated explicitly and droppable if it fails). **Hard requirement on the build +order: persist per-row `n`, `SE`, and pre-adjustment `p`** — without them R2/R4 can never be adjudicated. + +**Re-adjudication flagged:** p_win→CLV, the skew audit, **proj-v1.1's "NOT PROVEN" (judged against the +collapsed champion — not final)**, the C1 takeable floor, the calibration curves, and **ROI-by-grade — +with B/C inverted, "MLB-C +4.57%" is likely an artifact of a meaningless letter.** diff --git a/specs/STATE.md b/specs/STATE.md index e584ae5..f36e2ff 100644 --- a/specs/STATE.md +++ b/specs/STATE.md @@ -463,6 +463,61 @@ > `sports.{sport}` and reads ACCRUING until its own n≥20 clears; never `overall`, which would silently > borrow MLB/WNBA credibility. +> ## 🧪 FULL-OUTPUT GRADE MAPPING + COLLAPSE COST 2026-07-30 (report-only) → **`specs/full-output-grade-mapping.md`** +> Track-B 1 of 3. Nothing built, reconnected, or promoted. **PREMISE CORRECTED AGAIN: the three-layer +> engine is BUILT but NOT WIRED and NOT DEPLOYED** (re-verified independently — 0 python refs in every +> grade-path file, 0 python lines in `Dockerfile`; there is no `engine1Adapter`, the real files are +> `utils/gradeAdapter.js` + `analyzeViaEngine1.js`). So no posterior/CI/similarity prior is produced +> today; Phase 1 could not inventory a live three-layer output and Phase 3 could not diff against one. +> Phase 3 was run instead against the collapse that ACTUALLY exists — which is **worse** than the +> premise describes, and measurable now. +> **THREE COLLAPSES (VERIFIED), not one.** (A) `estimateProbability` returns +> `{p_over,p_under,components{base,recency,weighted,opp_adjustment,home_adjustment,consistency_adjustment,cv}}` +> and `analyzeViaEngine1:521-524` keeps ONLY the scalar — `est.components` is attached to nothing. +> **(B) THE SEVERE ONE: `p_win` never reaches the grade at all** — the letter is `engine1`'s factor +> index and `engine1.js` has ZERO probability references, so the probability isn't collapsed INTO the +> grade, it's excluded FROM it. (C) `grade_thresholds.json` (PROBABILITY→GRADE) is read BACKWARDS to +> manufacture `confidence`. **0.2: market-efficiency scaling is NEVER COMPUTED** — a gap, not a +> second collapse. +> **🔴 PHASE 3 — THE COLLAPSE IS COSTLY ON MLB, AND UN-COLLAPSING DOES NOT HELP WNBA.** 354 settled +> rows carrying both the served letter and the locked pre-game `p_win` (forward test, NOT lookahead). +> Grade→outcome point-biserial r: **champion letter 0.0050 (p≈0.93, NULL)** vs **probability letter +> 0.1313 (p≈0.013)**. Per sport: **MLB champ 0.0686 n.s. vs prob 0.2356 (p≈0.0004, n=224)** — +> **WNBA champ −0.0986 vs prob −0.1258 (n=130, BOTH INVERSE)**. The pooled number is MLB's signal +> diluted by WNBA's inversion; this independently corroborates the 07-26 calibration finding that the +> champion does not discriminate on WNBA. **→ the full-output challenger must be MLB-FIRST; shipping +> it for WNBA on the pooled number would ship an anti-predictive grade.** +> **THE SERVED LETTER IS INVERTED BETWEEN ITS ONLY TWO POPULATED TIERS: B hits 52.4% (n=168), C hits +> 56.9% (n=174).** A user reading B as better than C is reading noise. Probability letters spread +> 30.8% (D) → 70.0% (B+), use 10-11 of 11 letters vs the champion's 3-4, and split ROI **−1.42% +> (A-family, n=78) vs −26.62% (C-/D/F, n=51) — a 25-point spread**. Banding is NOT the lossy part +> (0.1313 banded vs 0.1349 raw). +> **PHASE 2 — five EXPLICIT, FALSIFIABLE rules specced** (none assumed to be an improvement): +> R1 posterior→letter via the table read FORWARD · **R2 uncertainty grades DOWN, stated not smuggled: +> `p_adj = 0.5 + (p−0.5)·(1 − k·min(1, 1.96·SE/W0))`, k=1, W0=0.15 — falsifiable: rows R2 moves down +> must hit closer to their NEW band or R2 is WRONG and gets dropped** · R3 archetype adjusts the +> PROJECTION never the letter (drop if MAE doesn't improve) · R4 efficiency scales the THRESHOLD per +> sport, `E_sport` FIT from each sport's own record — success = equal hit rate per letter ACROSS +> sports · R5 abstain below min instances, never a default C. +> **🔴 HARD REQUIREMENT ON THE NEXT ORDER: R2/R3/R4 are UNMEASURABLE today — per-row instance count +> `n` is NOT stored on the ledger.** The challenger build MUST emit and persist `n`, `SE`, and the +> pre-adjustment `p`, or the CI-width and efficiency rules can never be adjudicated. +> **3.7 CONSUMERS:** nothing consumes a distribution, so a distribution-based grade is safe IF it +> still emits a letter; what changes is the letter DISTRIBUTION — `tierGating`, byGrade buckets +> (AccuracyBadge/ModelRecord/TierRecord), `outcomeService.gradeBucket`, `isAB` in hero + +> deskShowcase (the hero pool grows), `GRADE_RANK`/`selectTopGrades`, `capper_minimum_grade:'A-'`, +> newsletter templates. **A-GRADES: probability grading emits 53 A-family rows where the champion +> emitted 2. This is NOT the forbidden "rescale to mint A's"** — it is a measurably more informative +> basis (new information), and the A's earn it (A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%) — **but A- +> hits 44.0%, breaking top-tier monotonicity, so the A-RATED marketing hold STAYS until the +> challenger's own forward record shows a monotone top tier.** +> **PHASE 4 RE-ADJUDICATION LIST (flagged, not re-run):** champion p_win→CLV 0.375 · the over-side +> skew audit · **proj-v1.1's "NOT PROVEN" — it was judged against the COLLAPSED champion, so its +> death is NOT final** · the C1 takeable-floor derivation · the 07-26 calibration curves (its WNBA +> finding is corroborated and looks robust; MLB needs re-running) · **ROI-by-grade — with B and C +> inverted, "MLB-C is the +4.57% profitable segment" is very likely an artifact of a meaningless +> letter, not a real segment** · `confidence` on every historical row (never use it as a weight). + - **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows`→`closing_captures` (append-only, provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221` + internal endpoint; `ledgerService.attachClosingProb`→`closing_prob` (de-vigs both raw sides, diff --git a/specs/full-output-grade-mapping.md b/specs/full-output-grade-mapping.md new file mode 100644 index 0000000..46f72db --- /dev/null +++ b/specs/full-output-grade-mapping.md @@ -0,0 +1,224 @@ +# SPEC — FULL-OUTPUT GRADE MAPPING + COLLAPSE-COST MEASUREMENT +Report-only, 2026-07-30. **Nothing built, reconnected, promoted, or changed.** +Track-B order 1 of 3. The challenger BUILD is the next order, gated on this. + +--- + +## REVIEW ZERO — the premise needs one correction before anything else + +**The order's premise says the three-layer engine is "BUILT and WIRED."** Re-verified +independently this order: **it is BUILT but NOT WIRED and NOT DEPLOYED.** Every grade-path +file (`gradeSlateService`, `analyzeViaEngine1`, `engine1`, `featureCache`, +`probabilityEstimator`) contains **zero** references to the Python service, and `Dockerfile` +contains **zero** python/pip/requirements lines. There is also no `engine1Adapter` — the real +files are `utils/gradeAdapter.js` and `intelligence/analyzeViaEngine1.js`. + +**Consequence for this order:** there is no Layer-1 similarity prior, no Layer-2 posterior, +and no confidence interval being produced per prop today. **Phase 1 cannot inventory a live +three-layer output, and Phase 3 cannot diff against one.** So Phase 3 was executed against +the collapse that *actually exists* — which turns out to be more severe than the premise +describes, and measurable right now. + +### 0.1 The collapse points — THREE, not one (VERIFIED) + +**Collapse A — the estimator's components are discarded.** +`probabilityEstimator.estimateProbability` returns +`{ p_over, p_under, components: { base, recency, weighted, opp_adjustment, home_adjustment, +consistency_adjustment, cv } }`. At `analyzeViaEngine1.js:521-524` only the scalar survives: + + const est = estimateProbability({ gameLogs: meta.gameLogs, line: prop.line, ... }); + const pWin = dir === 'under' ? (1 - est.p_over) : est.p_over; + if (pWin != null) legacy.p_win = Math.round(pWin * 1000) / 1000; + +`est.components` is never attached to anything. Discarded: the empirical base rate, the +recency rate, each adjustment's magnitude, and `cv` (the only uncertainty proxy produced). + +**Collapse B — THE SEVERE ONE: `p_win` never reaches the grade at all.** +The letter comes from `engine1.gradeProp` (`idx = NEUTRAL_INDEX(3); idx += f.delta`), and +`engine1.js` contains **zero** references to `p_win` or any probability. So the probability is +not "collapsed into the grade" — **it is computed, stored on the payload, and excluded from +grading entirely.** + +**Collapse C — `grade_thresholds.json` is read BACKWARDS.** The table maps +PROBABILITY → GRADE (`A+ 0.85-1.00`). The live path picks a letter from the factor index and +then reads that letter's band MIDPOINT to manufacture `confidence`. + +**NOT produced anywhere today (so not "discarded" — absent):** a posterior distribution, a +confidence interval, a similarity prior, an instance count. `cv` is a coefficient of variation +of the STAT, not a CI on the probability. + +### 0.2 Market-efficiency-per-sport — SPECCED-BUT-ABSENT, not flattened (VERIFIED) +It is not computed-then-flattened; it is **never computed**. One global 11-band scale with no +sport dimension. The spec's MLB 0.55 / NBA-stars 0.80 exists nowhere in code. So it is a +second *gap*, not a second collapse. + +--- + +## PHASE 1 — INVENTORY + +| Component | Produced today? | Reaches the grade? | +|---|---|---| +| Layer-1 similarity prior | **NO** (engine not deployed) | — | +| Instance count / min-instance fallback | **NO** | — | +| Layer-2 Bayesian posterior | **NO** | — | +| Posterior confidence interval | **NO** | — | +| Archetype context | YES (snapshot classify) | **NO** — display + challenger only | +| `p_over` point estimate | YES | **NO** (Collapse B) | +| Estimator components / `cv` | YES, then dropped | **NO** (Collapse A) | +| Factor deltas (l5/l20, opp_rank, rest, usage) | YES | **YES — these ARE the grade** | + +**Min-instance fallback (spec: 15): CANNOT DETERMINE / does not fire.** The similarity engine +never runs. The code's only abstention rule is `ABSTENTION_RULES.similar_games_below = 3` +(`bayesian.py:46`); the only `15` in the file is NBA `min_minutes_per_game`. When data is thin +today the live path does not fall back to a prior — `analyzeViaEngine1` **REFUSES** the read +(`insufficient_data: true`, grade null) if `projectionFor` finds no reference. + +--- + +## PHASE 2 — THE DISTRIBUTION → GRADE MAPPING (design, on paper) + +Five explicit rules. **Each is stated so it can be falsified on the forward instrument. None +of them is assumed to be an improvement.** + +**R1 — BASE: the posterior sets the letter.** +`grade = band(p_posterior)` using `grade_thresholds.json` read FORWARD (the direction it was +written for). *Testable:* within each band, realized hit rate should fall inside the band. + +**R2 — UNCERTAINTY GRADES DOWN (the explicit rule, not smuggled).** +Compute the posterior's SE; shrink toward 0.5 by the CI half-width before banding: + + halfWidth = z * SE(p) // z = 1.96 + p_adj = 0.5 + (p - 0.5) * (1 - k * min(1, halfWidth / W0)) + grade = band(p_adj) // defaults k = 1.0, W0 = 0.15 + +So two props at the same `P(>=line)` grade DIFFERENTLY when their uncertainty differs — the +wider one grades lower. *Testable, and falsifiable:* rows that R2 moves DOWN should hit closer +to their NEW band than their OLD band. If they hit closer to the old band, **R2 is wrong and +must be dropped** — high uncertainty would then be noise, not a reason to downgrade. + +**R3 — ARCHETYPE ADJUSTS THE PROJECTION, NEVER THE LETTER.** +Archetype modifies the projection *before* the probability is computed; it must never be a +post-hoc letter bump. *Testable:* archetype-adjusted projections should lower MAE against +actuals vs unadjusted. If MAE does not improve, R3 is dropped. + +**R4 — MARKET-EFFICIENCY SCALES THE THRESHOLD, NOT THE PROBABILITY.** + + threshold_sport(letter) = 0.5 + (threshold_base(letter) - 0.5) * E_sport + +Higher `E_sport` (more efficient market) demands more probability for the same letter. +*Testable, and this is the definition of the goal:* after scaling, **the realized hit rate for +a given letter should be EQUAL across sports.** That is what "an A means the same thing +everywhere" cashes out to, and it is measurable. `E_sport` is FIT from each sport's own +accrued record — never hand-set. + +**R5 — ABSTENTION, not a default C.** Below the min-instance threshold, refuse (grade null), +matching today's honest refusal behaviour. The threshold is a declared parameter (code says 3; +the spec's 15 is unreconciled) and must be set explicitly, not inherited by accident. + +--- + +## PHASE 3 — THE COLLAPSE COST, MEASURED + +**Method.** 354 SETTLED public ledger rows carrying both the served champion letter and the +stored `p_win`. `p_win` was computed pre-game and locked, so this is a forward test, **not +lookahead**. The champion letter is compared against the probability letter obtained by +reading `grade_thresholds.json` FORWARD (R1 only — R2/R3/R4 are not measurable from stored +data, see the gap note below). + +### Discrimination — grade vs realized outcome (point-biserial r) + +| population | n | **champion letter** | **probability letter** | raw `p_win` | +|---|---|---|---|---| +| **ALL** | 354 | **0.0050** (t≈0.09, p≈0.93 — null) | **0.1313** (t≈2.49, **p≈0.013**) | 0.1349 | +| **MLB** | 224 | 0.0686 (n.s.) | **0.2356** (t≈3.61, **p≈0.0004**) | — | +| **WNBA** | 130 | **−0.0986** | **−0.1258** | — | + +### Hit rate by letter + +| champion | n | hit% | | probability | n | hit% | +|---|---|---|---|---|---|---| +| A | 2 | 50.0 | | A+ | 24 | 62.5 | +| B | 168 | **52.4** | | A | 29 | 65.5 | +| C | 174 | **56.9** | | A- | 25 | 44.0 | +| D | 5 | 20.0 | | B+ | 30 | **70.0** | +| F | 5 | 20.0 | | B / B- | 54 / 48 | 53.7 / 56.3 | +| | | | | C+ / C / C- | 43 / 50 / 15 | 53.5 / 50.0 / 53.3 | +| | | | | D / F | 26 / 10 | **30.8** / 40.0 | + +**Flat-stake ROI:** probability-top (A-family, n=78) **−1.42%** vs probability-bottom +(C-/D/F, n=51) **−26.62%** — a **25-point spread**. The champion's top tier is n=2 (unusable). +Letters actually used: champion **3-4**; probability **10-11** of 11. + +### 3.6 — VERDICT: **COSTLY on MLB, and un-collapsing does NOT help WNBA** + +- **The champion's served letter is uninformative.** r = 0.005 overall, and it is **INVERTED + between its only two populated tiers — B hits 52.4% while C hits 56.9%.** A user reading B + as better than C is reading noise. +- **MLB: the collapse is costly and the cost is significant.** The discarded probability + carries real signal (r = 0.236, p ≈ 0.0004) where the served letter carries effectively none + (0.069, n.s.). **This is the strongest single finding of the model line to date.** +- **WNBA: un-collapsing makes it WORSE, not better** (prob r = **−0.126**, i.e. inverse). The + pooled r = 0.131 is MLB's signal diluted by WNBA's inversion. This independently corroborates + the 2026-07-26 calibration diagnosis ("champion does not discriminate on WNBA"; Brier worse + than always-0.5). +- **Therefore the full-output challenger must be MLB-FIRST.** Shipping it for WNBA on the + strength of a pooled number would ship an anti-predictive grade. WNBA needs a different fix. +- Banding costs almost nothing vs the raw probability (0.1313 vs 0.1349) — the 11-band scale is + not the lossy part. + +### GAP — R2/R3/R4 are NOT measured here (CANNOT DETERMINE) +`SE(p)` needs the per-row instance count `n`, which **is not stored on the ledger**. So the +CI-width rule, archetype-on-projection, and efficiency scaling are **specified but unmeasured**. +**The challenger build must emit and persist `n`, `SE`, and the pre-adjustment `p`** or R2/R4 +can never be adjudicated. That is a hard requirement on the next order. + +### 3.7 — Consumers that assume the grade is a POINT +Nothing consumes a distribution, so a distribution-based grade is safe **provided it still +emits a letter**. What DOES change is the letter DISTRIBUTION, and these depend on it: +`tierGating` (Desk gates), `AccuracyBadge` / `ModelRecord` / `TierRecord` byGrade buckets, +`outcomeService.gradeBucket` + `TIERS`, `heroPropService.isAB` and `deskShowcaseService.isAB` +(A/B filters — the hero pool would grow), `gradeRanking.GRADE_RANK`, `selectTopGrades`, +`grade_thresholds.json:capper_minimum_grade = 'A-'`, and the newsletter/media templates. + +**🔴 The A-grade question, stated precisely.** Probability grading would emit **53 A-family +rows where the champion emitted 2**. CLAUDE.md's permanent founder ruling forbids *rescaling +thresholds to mint A's*. **This is not that** — it is grading on a different and measurably more +informative basis (r 0.236 vs 0.069 on MLB), which is new information, and the A's earn it +(prob-A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%). **But one honest caveat: A- hits 44.0%, breaking +monotonicity at the top.** The A-RATED marketing hold should stay until the challenger's own +forward record shows a monotone top tier — the ruling's spirit (never sell a relabelled B as an +A) is satisfied by evidence, not by the mapping alone. + +--- + +## PHASE 4 — RE-ADJUDICATION LIST (collapsed-output verdicts, flagged not re-run) + +Every one of these was measured on the collapsed output and must be re-measured on the full +output before being trusted forward: + +1. **The champion's p_win→CLV edge** (partial r = 0.375, p ≈ 0.003, n = 62 takeable MLB overs) + — measured on the collapsed p_win. Direction is likely preserved (it already used p_win, not + the letter), but the magnitude is a collapsed-output number. +2. **The over-side skew audit** ("SURVIVES BASELINE", +7.14pt marginal) — same basis. +3. **proj-v1.1's "NOT PROVEN"** — it was judged against the COLLAPSED champion. A full-output + champion is a different, stronger benchmark, so **proj-v1.1's death is not final** — it must + be re-run against the new champion, and could fare better or worse. +4. **The takeable-floor derivation (C1)** — ROI-by-price buckets were computed on rows graded by + the collapsed engine; the population selection itself is collapsed-output-conditioned. +5. **The calibration diagnosis (2026-07-26)** — Brier/reliability/resolution on collapsed p_win. + Its WNBA finding is *corroborated* by this order and looks robust; its MLB numbers need re-running. +6. **ROI-by-grade (MLB-C +4.57%, MLB-B −1.02%, WNBA −5%)** — these are grade buckets from a + letter now measured at r ≈ 0.005. **Given B and C are inverted, "MLB-C is the profitable + segment" is very likely an artifact of a meaningless letter, not a real segment.** +7. **`confidence` on every historical row** — a band-midpoint of a letter with no predictive + content; it should not be used as a weight in any future analysis. + +--- + +## TAGS +VERIFIED: the three collapse points; the engine is not wired/deployed; discrimination numbers +(n=354, per sport); consumer list. CANNOT DETERMINE: R2/R3/R4 effects (per-row `n` not stored); +the spec's min-15 instance rule (absent from code). BLOCKED: none. + +## HELD +No grade change, no challenger build, no promotion, no reconnection, no retro-grading.