report: full-output grade mapping spec + collapse cost measured

Report-only. Nothing built, reconnected, or promoted.

Premise corrected again: the three-layer engine is BUILT but NOT WIRED and NOT
DEPLOYED (0 python refs in every grade-path file, 0 python in Dockerfile; there
is no engine1Adapter). So no posterior/CI/similarity prior exists to inventory
or diff. Measured against the collapse that actually exists instead.

THREE collapses, not one: (A) estimateProbability's components discarded at
analyzeViaEngine1:521-524; (B) THE SEVERE ONE - p_win never reaches the grade
at all (engine1.js has zero probability references), so the probability is
excluded from grading rather than collapsed into it; (C) grade_thresholds.json
(probability->grade) read backwards to manufacture confidence.
Market-efficiency scaling is never computed - a gap, not a collapse.

MEASURED on 354 settled rows carrying the served letter and the locked pre-game
p_win (forward, not lookahead). Grade->outcome point-biserial r: champion letter
0.0050 (p~0.93, null) vs probability letter 0.1313 (p~0.013). Per sport: MLB
champ 0.0686 n.s. vs prob 0.2356 (p~0.0004); WNBA champ -0.0986 vs prob -0.1258
- BOTH INVERSE. The served letter is inverted between its only two populated
tiers (B 52.4% n=168 vs C 56.9% n=174).

Verdict: costly on MLB, and un-collapsing does NOT help WNBA -> the challenger
must be MLB-FIRST. Five falsifiable mapping rules specced, incl. R2
(uncertainty grades down) stated explicitly and droppable if it fails.

Hard requirement on the next order: persist per-row n, SE and pre-adjustment p,
or R2/R4 can never be adjudicated (not stored today).

Re-adjudication list flagged incl. proj-v1.1's NOT PROVEN verdict (judged
against the collapsed champion, so not final) and ROI-by-grade (with B/C
inverted, the MLB-C +4.57% segment is likely an artifact of a meaningless
letter).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
Kev
2026-07-30 23:47:11 -04:00
parent dd98b0b614
commit 708f0fde5c
3 changed files with 309 additions and 0 deletions
+30
View File
@@ -426,3 +426,33 @@ are absent entirely.
Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting
Bayesian to Node) → wire Layer 2 → reconnect Layer 3 forward → Layer 1 → per-sport efficiency → Bayesian to Node) → wire Layer 2 → reconnect Layer 3 forward → Layer 1 → per-sport efficiency →
challenger promotion → sport-as-module → NFL/CFB. MLB is specced as the reference module. challenger promotion → sport-as-module → NFL/CFB. MLB is specced as the reference module.
---
# FULL-OUTPUT GRADE MAPPING + COLLAPSE COST — 2026-07-30 (report-only) → `specs/full-output-grade-mapping.md`
Track-B 1 of 3. The three-layer engine is **BUILT but NOT WIRED and NOT DEPLOYED** (re-verified), so no
posterior/CI is produced today. Measured instead against the collapse that actually exists — three of
them: estimator components dropped; **`p_win` excluded from the grade entirely** (the severe one — the
letter is a factor index with zero probability input); and `grade_thresholds.json` read backwards to
manufacture `confidence`. Market-efficiency scaling is never computed at all.
**THE MEASUREMENT (354 settled rows, locked pre-game p_win — forward, not lookahead):**
| basis | ALL (n=354) | MLB (n=224) | WNBA (n=130) |
|---|---|---|---|
| champion letter → outcome r | **0.0050** (p≈0.93, null) | 0.0686 (n.s.) | 0.0986 |
| probability letter → outcome r | **0.1313** (p≈0.013) | **0.2356** (p≈0.0004) | **0.1258** |
**The served letter is INVERTED between its only two populated tiers — B 52.4% (n=168) vs C 56.9%
(n=174).** Probability letters spread 30.8%→70.0%, use 10-11 of 11 letters (champion uses 3-4), and
split ROI 1.42% (A-family n=78) vs 26.62% (C-/D/F n=51).
**Verdict: costly on MLB, and un-collapsing does NOT help WNBA** (both correlations inverse there) —
so the full-output challenger must be **MLB-FIRST**. Five falsifiable mapping rules are specced (R2
"uncertainty grades down" stated explicitly and droppable if it fails). **Hard requirement on the build
order: persist per-row `n`, `SE`, and pre-adjustment `p`** — without them R2/R4 can never be adjudicated.
**Re-adjudication flagged:** p_win→CLV, the skew audit, **proj-v1.1's "NOT PROVEN" (judged against the
collapsed champion — not final)**, the C1 takeable floor, the calibration curves, and **ROI-by-grade —
with B/C inverted, "MLB-C +4.57%" is likely an artifact of a meaningless letter.**
+55
View File
@@ -463,6 +463,61 @@
> `sports.{sport}` and reads ACCRUING until its own n≥20 clears; never `overall`, which would silently > `sports.{sport}` and reads ACCRUING until its own n≥20 clears; never `overall`, which would silently
> borrow MLB/WNBA credibility. > borrow MLB/WNBA credibility.
> ## 🧪 FULL-OUTPUT GRADE MAPPING + COLLAPSE COST 2026-07-30 (report-only) → **`specs/full-output-grade-mapping.md`**
> Track-B 1 of 3. Nothing built, reconnected, or promoted. **PREMISE CORRECTED AGAIN: the three-layer
> engine is BUILT but NOT WIRED and NOT DEPLOYED** (re-verified independently — 0 python refs in every
> grade-path file, 0 python lines in `Dockerfile`; there is no `engine1Adapter`, the real files are
> `utils/gradeAdapter.js` + `analyzeViaEngine1.js`). So no posterior/CI/similarity prior is produced
> today; Phase 1 could not inventory a live three-layer output and Phase 3 could not diff against one.
> Phase 3 was run instead against the collapse that ACTUALLY exists — which is **worse** than the
> premise describes, and measurable now.
> **THREE COLLAPSES (VERIFIED), not one.** (A) `estimateProbability` returns
> `{p_over,p_under,components{base,recency,weighted,opp_adjustment,home_adjustment,consistency_adjustment,cv}}`
> and `analyzeViaEngine1:521-524` keeps ONLY the scalar — `est.components` is attached to nothing.
> **(B) THE SEVERE ONE: `p_win` never reaches the grade at all** — the letter is `engine1`'s factor
> index and `engine1.js` has ZERO probability references, so the probability isn't collapsed INTO the
> grade, it's excluded FROM it. (C) `grade_thresholds.json` (PROBABILITY→GRADE) is read BACKWARDS to
> manufacture `confidence`. **0.2: market-efficiency scaling is NEVER COMPUTED** — a gap, not a
> second collapse.
> **🔴 PHASE 3 — THE COLLAPSE IS COSTLY ON MLB, AND UN-COLLAPSING DOES NOT HELP WNBA.** 354 settled
> rows carrying both the served letter and the locked pre-game `p_win` (forward test, NOT lookahead).
> Grade→outcome point-biserial r: **champion letter 0.0050 (p≈0.93, NULL)** vs **probability letter
> 0.1313 (p≈0.013)**. Per sport: **MLB champ 0.0686 n.s. vs prob 0.2356 (p≈0.0004, n=224)** —
> **WNBA champ 0.0986 vs prob 0.1258 (n=130, BOTH INVERSE)**. The pooled number is MLB's signal
> diluted by WNBA's inversion; this independently corroborates the 07-26 calibration finding that the
> champion does not discriminate on WNBA. **→ the full-output challenger must be MLB-FIRST; shipping
> it for WNBA on the pooled number would ship an anti-predictive grade.**
> **THE SERVED LETTER IS INVERTED BETWEEN ITS ONLY TWO POPULATED TIERS: B hits 52.4% (n=168), C hits
> 56.9% (n=174).** A user reading B as better than C is reading noise. Probability letters spread
> 30.8% (D) → 70.0% (B+), use 10-11 of 11 letters vs the champion's 3-4, and split ROI **1.42%
> (A-family, n=78) vs 26.62% (C-/D/F, n=51) — a 25-point spread**. Banding is NOT the lossy part
> (0.1313 banded vs 0.1349 raw).
> **PHASE 2 — five EXPLICIT, FALSIFIABLE rules specced** (none assumed to be an improvement):
> R1 posterior→letter via the table read FORWARD · **R2 uncertainty grades DOWN, stated not smuggled:
> `p_adj = 0.5 + (p0.5)·(1 k·min(1, 1.96·SE/W0))`, k=1, W0=0.15 — falsifiable: rows R2 moves down
> must hit closer to their NEW band or R2 is WRONG and gets dropped** · R3 archetype adjusts the
> PROJECTION never the letter (drop if MAE doesn't improve) · R4 efficiency scales the THRESHOLD per
> sport, `E_sport` FIT from each sport's own record — success = equal hit rate per letter ACROSS
> sports · R5 abstain below min instances, never a default C.
> **🔴 HARD REQUIREMENT ON THE NEXT ORDER: R2/R3/R4 are UNMEASURABLE today — per-row instance count
> `n` is NOT stored on the ledger.** The challenger build MUST emit and persist `n`, `SE`, and the
> pre-adjustment `p`, or the CI-width and efficiency rules can never be adjudicated.
> **3.7 CONSUMERS:** nothing consumes a distribution, so a distribution-based grade is safe IF it
> still emits a letter; what changes is the letter DISTRIBUTION — `tierGating`, byGrade buckets
> (AccuracyBadge/ModelRecord/TierRecord), `outcomeService.gradeBucket`, `isAB` in hero +
> deskShowcase (the hero pool grows), `GRADE_RANK`/`selectTopGrades`, `capper_minimum_grade:'A-'`,
> newsletter templates. **A-GRADES: probability grading emits 53 A-family rows where the champion
> emitted 2. This is NOT the forbidden "rescale to mint A's"** — it is a measurably more informative
> basis (new information), and the A's earn it (A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%) — **but A-
> hits 44.0%, breaking top-tier monotonicity, so the A-RATED marketing hold STAYS until the
> challenger's own forward record shows a monotone top tier.**
> **PHASE 4 RE-ADJUDICATION LIST (flagged, not re-run):** champion p_win→CLV 0.375 · the over-side
> skew audit · **proj-v1.1's "NOT PROVEN" — it was judged against the COLLAPSED champion, so its
> death is NOT final** · the C1 takeable-floor derivation · the 07-26 calibration curves (its WNBA
> finding is corroborated and looks robust; MLB needs re-running) · **ROI-by-grade — with B and C
> inverted, "MLB-C is the +4.57% profitable segment" is very likely an artifact of a meaningless
> letter, not a real segment** · `confidence` on every historical row (never use it as a weight).
- **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows``closing_captures` (append-only, - **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows``closing_captures` (append-only,
provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221` provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221`
+ internal endpoint; `ledgerService.attachClosingProb``closing_prob` (de-vigs both raw sides, + internal endpoint; `ledgerService.attachClosingProb``closing_prob` (de-vigs both raw sides,
+224
View File
@@ -0,0 +1,224 @@
# SPEC — FULL-OUTPUT GRADE MAPPING + COLLAPSE-COST MEASUREMENT
Report-only, 2026-07-30. **Nothing built, reconnected, promoted, or changed.**
Track-B order 1 of 3. The challenger BUILD is the next order, gated on this.
---
## REVIEW ZERO — the premise needs one correction before anything else
**The order's premise says the three-layer engine is "BUILT and WIRED."** Re-verified
independently this order: **it is BUILT but NOT WIRED and NOT DEPLOYED.** Every grade-path
file (`gradeSlateService`, `analyzeViaEngine1`, `engine1`, `featureCache`,
`probabilityEstimator`) contains **zero** references to the Python service, and `Dockerfile`
contains **zero** python/pip/requirements lines. There is also no `engine1Adapter` — the real
files are `utils/gradeAdapter.js` and `intelligence/analyzeViaEngine1.js`.
**Consequence for this order:** there is no Layer-1 similarity prior, no Layer-2 posterior,
and no confidence interval being produced per prop today. **Phase 1 cannot inventory a live
three-layer output, and Phase 3 cannot diff against one.** So Phase 3 was executed against
the collapse that *actually exists* — which turns out to be more severe than the premise
describes, and measurable right now.
### 0.1 The collapse points — THREE, not one (VERIFIED)
**Collapse A — the estimator's components are discarded.**
`probabilityEstimator.estimateProbability` returns
`{ p_over, p_under, components: { base, recency, weighted, opp_adjustment, home_adjustment,
consistency_adjustment, cv } }`. At `analyzeViaEngine1.js:521-524` only the scalar survives:
const est = estimateProbability({ gameLogs: meta.gameLogs, line: prop.line, ... });
const pWin = dir === 'under' ? (1 - est.p_over) : est.p_over;
if (pWin != null) legacy.p_win = Math.round(pWin * 1000) / 1000;
`est.components` is never attached to anything. Discarded: the empirical base rate, the
recency rate, each adjustment's magnitude, and `cv` (the only uncertainty proxy produced).
**Collapse B — THE SEVERE ONE: `p_win` never reaches the grade at all.**
The letter comes from `engine1.gradeProp` (`idx = NEUTRAL_INDEX(3); idx += f.delta`), and
`engine1.js` contains **zero** references to `p_win` or any probability. So the probability is
not "collapsed into the grade" — **it is computed, stored on the payload, and excluded from
grading entirely.**
**Collapse C — `grade_thresholds.json` is read BACKWARDS.** The table maps
PROBABILITY → GRADE (`A+ 0.85-1.00`). The live path picks a letter from the factor index and
then reads that letter's band MIDPOINT to manufacture `confidence`.
**NOT produced anywhere today (so not "discarded" — absent):** a posterior distribution, a
confidence interval, a similarity prior, an instance count. `cv` is a coefficient of variation
of the STAT, not a CI on the probability.
### 0.2 Market-efficiency-per-sport — SPECCED-BUT-ABSENT, not flattened (VERIFIED)
It is not computed-then-flattened; it is **never computed**. One global 11-band scale with no
sport dimension. The spec's MLB 0.55 / NBA-stars 0.80 exists nowhere in code. So it is a
second *gap*, not a second collapse.
---
## PHASE 1 — INVENTORY
| Component | Produced today? | Reaches the grade? |
|---|---|---|
| Layer-1 similarity prior | **NO** (engine not deployed) | — |
| Instance count / min-instance fallback | **NO** | — |
| Layer-2 Bayesian posterior | **NO** | — |
| Posterior confidence interval | **NO** | — |
| Archetype context | YES (snapshot classify) | **NO** — display + challenger only |
| `p_over` point estimate | YES | **NO** (Collapse B) |
| Estimator components / `cv` | YES, then dropped | **NO** (Collapse A) |
| Factor deltas (l5/l20, opp_rank, rest, usage) | YES | **YES — these ARE the grade** |
**Min-instance fallback (spec: 15): CANNOT DETERMINE / does not fire.** The similarity engine
never runs. The code's only abstention rule is `ABSTENTION_RULES.similar_games_below = 3`
(`bayesian.py:46`); the only `15` in the file is NBA `min_minutes_per_game`. When data is thin
today the live path does not fall back to a prior — `analyzeViaEngine1` **REFUSES** the read
(`insufficient_data: true`, grade null) if `projectionFor` finds no reference.
---
## PHASE 2 — THE DISTRIBUTION → GRADE MAPPING (design, on paper)
Five explicit rules. **Each is stated so it can be falsified on the forward instrument. None
of them is assumed to be an improvement.**
**R1 — BASE: the posterior sets the letter.**
`grade = band(p_posterior)` using `grade_thresholds.json` read FORWARD (the direction it was
written for). *Testable:* within each band, realized hit rate should fall inside the band.
**R2 — UNCERTAINTY GRADES DOWN (the explicit rule, not smuggled).**
Compute the posterior's SE; shrink toward 0.5 by the CI half-width before banding:
halfWidth = z * SE(p) // z = 1.96
p_adj = 0.5 + (p - 0.5) * (1 - k * min(1, halfWidth / W0))
grade = band(p_adj) // defaults k = 1.0, W0 = 0.15
So two props at the same `P(>=line)` grade DIFFERENTLY when their uncertainty differs — the
wider one grades lower. *Testable, and falsifiable:* rows that R2 moves DOWN should hit closer
to their NEW band than their OLD band. If they hit closer to the old band, **R2 is wrong and
must be dropped** — high uncertainty would then be noise, not a reason to downgrade.
**R3 — ARCHETYPE ADJUSTS THE PROJECTION, NEVER THE LETTER.**
Archetype modifies the projection *before* the probability is computed; it must never be a
post-hoc letter bump. *Testable:* archetype-adjusted projections should lower MAE against
actuals vs unadjusted. If MAE does not improve, R3 is dropped.
**R4 — MARKET-EFFICIENCY SCALES THE THRESHOLD, NOT THE PROBABILITY.**
threshold_sport(letter) = 0.5 + (threshold_base(letter) - 0.5) * E_sport
Higher `E_sport` (more efficient market) demands more probability for the same letter.
*Testable, and this is the definition of the goal:* after scaling, **the realized hit rate for
a given letter should be EQUAL across sports.** That is what "an A means the same thing
everywhere" cashes out to, and it is measurable. `E_sport` is FIT from each sport's own
accrued record — never hand-set.
**R5 — ABSTENTION, not a default C.** Below the min-instance threshold, refuse (grade null),
matching today's honest refusal behaviour. The threshold is a declared parameter (code says 3;
the spec's 15 is unreconciled) and must be set explicitly, not inherited by accident.
---
## PHASE 3 — THE COLLAPSE COST, MEASURED
**Method.** 354 SETTLED public ledger rows carrying both the served champion letter and the
stored `p_win`. `p_win` was computed pre-game and locked, so this is a forward test, **not
lookahead**. The champion letter is compared against the probability letter obtained by
reading `grade_thresholds.json` FORWARD (R1 only — R2/R3/R4 are not measurable from stored
data, see the gap note below).
### Discrimination — grade vs realized outcome (point-biserial r)
| population | n | **champion letter** | **probability letter** | raw `p_win` |
|---|---|---|---|---|
| **ALL** | 354 | **0.0050** (t≈0.09, p≈0.93 — null) | **0.1313** (t≈2.49, **p≈0.013**) | 0.1349 |
| **MLB** | 224 | 0.0686 (n.s.) | **0.2356** (t≈3.61, **p≈0.0004**) | — |
| **WNBA** | 130 | **0.0986** | **0.1258** | — |
### Hit rate by letter
| champion | n | hit% | | probability | n | hit% |
|---|---|---|---|---|---|---|
| A | 2 | 50.0 | | A+ | 24 | 62.5 |
| B | 168 | **52.4** | | A | 29 | 65.5 |
| C | 174 | **56.9** | | A- | 25 | 44.0 |
| D | 5 | 20.0 | | B+ | 30 | **70.0** |
| F | 5 | 20.0 | | B / B- | 54 / 48 | 53.7 / 56.3 |
| | | | | C+ / C / C- | 43 / 50 / 15 | 53.5 / 50.0 / 53.3 |
| | | | | D / F | 26 / 10 | **30.8** / 40.0 |
**Flat-stake ROI:** probability-top (A-family, n=78) **1.42%** vs probability-bottom
(C-/D/F, n=51) **26.62%** — a **25-point spread**. The champion's top tier is n=2 (unusable).
Letters actually used: champion **3-4**; probability **10-11** of 11.
### 3.6 — VERDICT: **COSTLY on MLB, and un-collapsing does NOT help WNBA**
- **The champion's served letter is uninformative.** r = 0.005 overall, and it is **INVERTED
between its only two populated tiers — B hits 52.4% while C hits 56.9%.** A user reading B
as better than C is reading noise.
- **MLB: the collapse is costly and the cost is significant.** The discarded probability
carries real signal (r = 0.236, p ≈ 0.0004) where the served letter carries effectively none
(0.069, n.s.). **This is the strongest single finding of the model line to date.**
- **WNBA: un-collapsing makes it WORSE, not better** (prob r = **0.126**, i.e. inverse). The
pooled r = 0.131 is MLB's signal diluted by WNBA's inversion. This independently corroborates
the 2026-07-26 calibration diagnosis ("champion does not discriminate on WNBA"; Brier worse
than always-0.5).
- **Therefore the full-output challenger must be MLB-FIRST.** Shipping it for WNBA on the
strength of a pooled number would ship an anti-predictive grade. WNBA needs a different fix.
- Banding costs almost nothing vs the raw probability (0.1313 vs 0.1349) — the 11-band scale is
not the lossy part.
### GAP — R2/R3/R4 are NOT measured here (CANNOT DETERMINE)
`SE(p)` needs the per-row instance count `n`, which **is not stored on the ledger**. So the
CI-width rule, archetype-on-projection, and efficiency scaling are **specified but unmeasured**.
**The challenger build must emit and persist `n`, `SE`, and the pre-adjustment `p`** or R2/R4
can never be adjudicated. That is a hard requirement on the next order.
### 3.7 — Consumers that assume the grade is a POINT
Nothing consumes a distribution, so a distribution-based grade is safe **provided it still
emits a letter**. What DOES change is the letter DISTRIBUTION, and these depend on it:
`tierGating` (Desk gates), `AccuracyBadge` / `ModelRecord` / `TierRecord` byGrade buckets,
`outcomeService.gradeBucket` + `TIERS`, `heroPropService.isAB` and `deskShowcaseService.isAB`
(A/B filters — the hero pool would grow), `gradeRanking.GRADE_RANK`, `selectTopGrades`,
`grade_thresholds.json:capper_minimum_grade = 'A-'`, and the newsletter/media templates.
**🔴 The A-grade question, stated precisely.** Probability grading would emit **53 A-family
rows where the champion emitted 2**. CLAUDE.md's permanent founder ruling forbids *rescaling
thresholds to mint A's*. **This is not that** — it is grading on a different and measurably more
informative basis (r 0.236 vs 0.069 on MLB), which is new information, and the A's earn it
(prob-A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%). **But one honest caveat: A- hits 44.0%, breaking
monotonicity at the top.** The A-RATED marketing hold should stay until the challenger's own
forward record shows a monotone top tier — the ruling's spirit (never sell a relabelled B as an
A) is satisfied by evidence, not by the mapping alone.
---
## PHASE 4 — RE-ADJUDICATION LIST (collapsed-output verdicts, flagged not re-run)
Every one of these was measured on the collapsed output and must be re-measured on the full
output before being trusted forward:
1. **The champion's p_win→CLV edge** (partial r = 0.375, p ≈ 0.003, n = 62 takeable MLB overs)
— measured on the collapsed p_win. Direction is likely preserved (it already used p_win, not
the letter), but the magnitude is a collapsed-output number.
2. **The over-side skew audit** ("SURVIVES BASELINE", +7.14pt marginal) — same basis.
3. **proj-v1.1's "NOT PROVEN"** — it was judged against the COLLAPSED champion. A full-output
champion is a different, stronger benchmark, so **proj-v1.1's death is not final** — it must
be re-run against the new champion, and could fare better or worse.
4. **The takeable-floor derivation (C1)** — ROI-by-price buckets were computed on rows graded by
the collapsed engine; the population selection itself is collapsed-output-conditioned.
5. **The calibration diagnosis (2026-07-26)** — Brier/reliability/resolution on collapsed p_win.
Its WNBA finding is *corroborated* by this order and looks robust; its MLB numbers need re-running.
6. **ROI-by-grade (MLB-C +4.57%, MLB-B 1.02%, WNBA 5%)** — these are grade buckets from a
letter now measured at r ≈ 0.005. **Given B and C are inverted, "MLB-C is the profitable
segment" is very likely an artifact of a meaningless letter, not a real segment.**
7. **`confidence` on every historical row** — a band-midpoint of a letter with no predictive
content; it should not be used as a weight in any future analysis.
---
## TAGS
VERIFIED: the three collapse points; the engine is not wired/deployed; discrimination numbers
(n=354, per sport); consumer list. CANNOT DETERMINE: R2/R3/R4 effects (per-row `n` not stored);
the spec's min-15 instance rule (absent from code). BLOCKED: none.
## HELD
No grade change, no challenger build, no promotion, no reconnection, no retro-grading.