report: full-output grade mapping spec + collapse cost measured
Report-only. Nothing built, reconnected, or promoted. Premise corrected again: the three-layer engine is BUILT but NOT WIRED and NOT DEPLOYED (0 python refs in every grade-path file, 0 python in Dockerfile; there is no engine1Adapter). So no posterior/CI/similarity prior exists to inventory or diff. Measured against the collapse that actually exists instead. THREE collapses, not one: (A) estimateProbability's components discarded at analyzeViaEngine1:521-524; (B) THE SEVERE ONE - p_win never reaches the grade at all (engine1.js has zero probability references), so the probability is excluded from grading rather than collapsed into it; (C) grade_thresholds.json (probability->grade) read backwards to manufacture confidence. Market-efficiency scaling is never computed - a gap, not a collapse. MEASURED on 354 settled rows carrying the served letter and the locked pre-game p_win (forward, not lookahead). Grade->outcome point-biserial r: champion letter 0.0050 (p~0.93, null) vs probability letter 0.1313 (p~0.013). Per sport: MLB champ 0.0686 n.s. vs prob 0.2356 (p~0.0004); WNBA champ -0.0986 vs prob -0.1258 - BOTH INVERSE. The served letter is inverted between its only two populated tiers (B 52.4% n=168 vs C 56.9% n=174). Verdict: costly on MLB, and un-collapsing does NOT help WNBA -> the challenger must be MLB-FIRST. Five falsifiable mapping rules specced, incl. R2 (uncertainty grades down) stated explicitly and droppable if it fails. Hard requirement on the next order: persist per-row n, SE and pre-adjustment p, or R2/R4 can never be adjudicated (not stored today). Re-adjudication list flagged incl. proj-v1.1's NOT PROVEN verdict (judged against the collapsed champion, so not final) and ROI-by-grade (with B/C inverted, the MLB-C +4.57% segment is likely an artifact of a meaningless letter). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc
This commit is contained in:
@@ -426,3 +426,33 @@ are absent entirely.
|
|||||||
Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting
|
Recovery is dependency-ordered in the map: decide the grading basis → pick a runtime (recommend porting
|
||||||
Bayesian to Node) → wire Layer 2 → reconnect Layer 3 forward → Layer 1 → per-sport efficiency →
|
Bayesian to Node) → wire Layer 2 → reconnect Layer 3 forward → Layer 1 → per-sport efficiency →
|
||||||
challenger promotion → sport-as-module → NFL/CFB. MLB is specced as the reference module.
|
challenger promotion → sport-as-module → NFL/CFB. MLB is specced as the reference module.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
# FULL-OUTPUT GRADE MAPPING + COLLAPSE COST — 2026-07-30 (report-only) → `specs/full-output-grade-mapping.md`
|
||||||
|
|
||||||
|
Track-B 1 of 3. The three-layer engine is **BUILT but NOT WIRED and NOT DEPLOYED** (re-verified), so no
|
||||||
|
posterior/CI is produced today. Measured instead against the collapse that actually exists — three of
|
||||||
|
them: estimator components dropped; **`p_win` excluded from the grade entirely** (the severe one — the
|
||||||
|
letter is a factor index with zero probability input); and `grade_thresholds.json` read backwards to
|
||||||
|
manufacture `confidence`. Market-efficiency scaling is never computed at all.
|
||||||
|
|
||||||
|
**THE MEASUREMENT (354 settled rows, locked pre-game p_win — forward, not lookahead):**
|
||||||
|
|
||||||
|
| basis | ALL (n=354) | MLB (n=224) | WNBA (n=130) |
|
||||||
|
|---|---|---|---|
|
||||||
|
| champion letter → outcome r | **0.0050** (p≈0.93, null) | 0.0686 (n.s.) | −0.0986 |
|
||||||
|
| probability letter → outcome r | **0.1313** (p≈0.013) | **0.2356** (p≈0.0004) | **−0.1258** |
|
||||||
|
|
||||||
|
**The served letter is INVERTED between its only two populated tiers — B 52.4% (n=168) vs C 56.9%
|
||||||
|
(n=174).** Probability letters spread 30.8%→70.0%, use 10-11 of 11 letters (champion uses 3-4), and
|
||||||
|
split ROI −1.42% (A-family n=78) vs −26.62% (C-/D/F n=51).
|
||||||
|
|
||||||
|
**Verdict: costly on MLB, and un-collapsing does NOT help WNBA** (both correlations inverse there) —
|
||||||
|
so the full-output challenger must be **MLB-FIRST**. Five falsifiable mapping rules are specced (R2
|
||||||
|
"uncertainty grades down" stated explicitly and droppable if it fails). **Hard requirement on the build
|
||||||
|
order: persist per-row `n`, `SE`, and pre-adjustment `p`** — without them R2/R4 can never be adjudicated.
|
||||||
|
|
||||||
|
**Re-adjudication flagged:** p_win→CLV, the skew audit, **proj-v1.1's "NOT PROVEN" (judged against the
|
||||||
|
collapsed champion — not final)**, the C1 takeable floor, the calibration curves, and **ROI-by-grade —
|
||||||
|
with B/C inverted, "MLB-C +4.57%" is likely an artifact of a meaningless letter.**
|
||||||
|
|||||||
@@ -463,6 +463,61 @@
|
|||||||
> `sports.{sport}` and reads ACCRUING until its own n≥20 clears; never `overall`, which would silently
|
> `sports.{sport}` and reads ACCRUING until its own n≥20 clears; never `overall`, which would silently
|
||||||
> borrow MLB/WNBA credibility.
|
> borrow MLB/WNBA credibility.
|
||||||
|
|
||||||
|
> ## 🧪 FULL-OUTPUT GRADE MAPPING + COLLAPSE COST 2026-07-30 (report-only) → **`specs/full-output-grade-mapping.md`**
|
||||||
|
> Track-B 1 of 3. Nothing built, reconnected, or promoted. **PREMISE CORRECTED AGAIN: the three-layer
|
||||||
|
> engine is BUILT but NOT WIRED and NOT DEPLOYED** (re-verified independently — 0 python refs in every
|
||||||
|
> grade-path file, 0 python lines in `Dockerfile`; there is no `engine1Adapter`, the real files are
|
||||||
|
> `utils/gradeAdapter.js` + `analyzeViaEngine1.js`). So no posterior/CI/similarity prior is produced
|
||||||
|
> today; Phase 1 could not inventory a live three-layer output and Phase 3 could not diff against one.
|
||||||
|
> Phase 3 was run instead against the collapse that ACTUALLY exists — which is **worse** than the
|
||||||
|
> premise describes, and measurable now.
|
||||||
|
> **THREE COLLAPSES (VERIFIED), not one.** (A) `estimateProbability` returns
|
||||||
|
> `{p_over,p_under,components{base,recency,weighted,opp_adjustment,home_adjustment,consistency_adjustment,cv}}`
|
||||||
|
> and `analyzeViaEngine1:521-524` keeps ONLY the scalar — `est.components` is attached to nothing.
|
||||||
|
> **(B) THE SEVERE ONE: `p_win` never reaches the grade at all** — the letter is `engine1`'s factor
|
||||||
|
> index and `engine1.js` has ZERO probability references, so the probability isn't collapsed INTO the
|
||||||
|
> grade, it's excluded FROM it. (C) `grade_thresholds.json` (PROBABILITY→GRADE) is read BACKWARDS to
|
||||||
|
> manufacture `confidence`. **0.2: market-efficiency scaling is NEVER COMPUTED** — a gap, not a
|
||||||
|
> second collapse.
|
||||||
|
> **🔴 PHASE 3 — THE COLLAPSE IS COSTLY ON MLB, AND UN-COLLAPSING DOES NOT HELP WNBA.** 354 settled
|
||||||
|
> rows carrying both the served letter and the locked pre-game `p_win` (forward test, NOT lookahead).
|
||||||
|
> Grade→outcome point-biserial r: **champion letter 0.0050 (p≈0.93, NULL)** vs **probability letter
|
||||||
|
> 0.1313 (p≈0.013)**. Per sport: **MLB champ 0.0686 n.s. vs prob 0.2356 (p≈0.0004, n=224)** —
|
||||||
|
> **WNBA champ −0.0986 vs prob −0.1258 (n=130, BOTH INVERSE)**. The pooled number is MLB's signal
|
||||||
|
> diluted by WNBA's inversion; this independently corroborates the 07-26 calibration finding that the
|
||||||
|
> champion does not discriminate on WNBA. **→ the full-output challenger must be MLB-FIRST; shipping
|
||||||
|
> it for WNBA on the pooled number would ship an anti-predictive grade.**
|
||||||
|
> **THE SERVED LETTER IS INVERTED BETWEEN ITS ONLY TWO POPULATED TIERS: B hits 52.4% (n=168), C hits
|
||||||
|
> 56.9% (n=174).** A user reading B as better than C is reading noise. Probability letters spread
|
||||||
|
> 30.8% (D) → 70.0% (B+), use 10-11 of 11 letters vs the champion's 3-4, and split ROI **−1.42%
|
||||||
|
> (A-family, n=78) vs −26.62% (C-/D/F, n=51) — a 25-point spread**. Banding is NOT the lossy part
|
||||||
|
> (0.1313 banded vs 0.1349 raw).
|
||||||
|
> **PHASE 2 — five EXPLICIT, FALSIFIABLE rules specced** (none assumed to be an improvement):
|
||||||
|
> R1 posterior→letter via the table read FORWARD · **R2 uncertainty grades DOWN, stated not smuggled:
|
||||||
|
> `p_adj = 0.5 + (p−0.5)·(1 − k·min(1, 1.96·SE/W0))`, k=1, W0=0.15 — falsifiable: rows R2 moves down
|
||||||
|
> must hit closer to their NEW band or R2 is WRONG and gets dropped** · R3 archetype adjusts the
|
||||||
|
> PROJECTION never the letter (drop if MAE doesn't improve) · R4 efficiency scales the THRESHOLD per
|
||||||
|
> sport, `E_sport` FIT from each sport's own record — success = equal hit rate per letter ACROSS
|
||||||
|
> sports · R5 abstain below min instances, never a default C.
|
||||||
|
> **🔴 HARD REQUIREMENT ON THE NEXT ORDER: R2/R3/R4 are UNMEASURABLE today — per-row instance count
|
||||||
|
> `n` is NOT stored on the ledger.** The challenger build MUST emit and persist `n`, `SE`, and the
|
||||||
|
> pre-adjustment `p`, or the CI-width and efficiency rules can never be adjudicated.
|
||||||
|
> **3.7 CONSUMERS:** nothing consumes a distribution, so a distribution-based grade is safe IF it
|
||||||
|
> still emits a letter; what changes is the letter DISTRIBUTION — `tierGating`, byGrade buckets
|
||||||
|
> (AccuracyBadge/ModelRecord/TierRecord), `outcomeService.gradeBucket`, `isAB` in hero +
|
||||||
|
> deskShowcase (the hero pool grows), `GRADE_RANK`/`selectTopGrades`, `capper_minimum_grade:'A-'`,
|
||||||
|
> newsletter templates. **A-GRADES: probability grading emits 53 A-family rows where the champion
|
||||||
|
> emitted 2. This is NOT the forbidden "rescale to mint A's"** — it is a measurably more informative
|
||||||
|
> basis (new information), and the A's earn it (A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%) — **but A-
|
||||||
|
> hits 44.0%, breaking top-tier monotonicity, so the A-RATED marketing hold STAYS until the
|
||||||
|
> challenger's own forward record shows a monotone top tier.**
|
||||||
|
> **PHASE 4 RE-ADJUDICATION LIST (flagged, not re-run):** champion p_win→CLV 0.375 · the over-side
|
||||||
|
> skew audit · **proj-v1.1's "NOT PROVEN" — it was judged against the COLLAPSED champion, so its
|
||||||
|
> death is NOT final** · the C1 takeable-floor derivation · the 07-26 calibration curves (its WNBA
|
||||||
|
> finding is corroborated and looks robust; MLB needs re-running) · **ROI-by-grade — with B and C
|
||||||
|
> inverted, "MLB-C is the +4.57% profitable segment" is very likely an artifact of a meaningless
|
||||||
|
> letter, not a real segment** · `confidence` on every historical row (never use it as a weight).
|
||||||
|
|
||||||
- **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows`→`closing_captures` (append-only,
|
- **Redirect EXISTS + WIRED:** `closingCapture.buildCaptureRows`→`closing_captures` (append-only,
|
||||||
provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221`
|
provenance: captured_at/book/line_type/both-prices/missed_reason) via `intradayRefreshService:221`
|
||||||
+ internal endpoint; `ledgerService.attachClosingProb`→`closing_prob` (de-vigs both raw sides,
|
+ internal endpoint; `ledgerService.attachClosingProb`→`closing_prob` (de-vigs both raw sides,
|
||||||
|
|||||||
@@ -0,0 +1,224 @@
|
|||||||
|
# SPEC — FULL-OUTPUT GRADE MAPPING + COLLAPSE-COST MEASUREMENT
|
||||||
|
Report-only, 2026-07-30. **Nothing built, reconnected, promoted, or changed.**
|
||||||
|
Track-B order 1 of 3. The challenger BUILD is the next order, gated on this.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## REVIEW ZERO — the premise needs one correction before anything else
|
||||||
|
|
||||||
|
**The order's premise says the three-layer engine is "BUILT and WIRED."** Re-verified
|
||||||
|
independently this order: **it is BUILT but NOT WIRED and NOT DEPLOYED.** Every grade-path
|
||||||
|
file (`gradeSlateService`, `analyzeViaEngine1`, `engine1`, `featureCache`,
|
||||||
|
`probabilityEstimator`) contains **zero** references to the Python service, and `Dockerfile`
|
||||||
|
contains **zero** python/pip/requirements lines. There is also no `engine1Adapter` — the real
|
||||||
|
files are `utils/gradeAdapter.js` and `intelligence/analyzeViaEngine1.js`.
|
||||||
|
|
||||||
|
**Consequence for this order:** there is no Layer-1 similarity prior, no Layer-2 posterior,
|
||||||
|
and no confidence interval being produced per prop today. **Phase 1 cannot inventory a live
|
||||||
|
three-layer output, and Phase 3 cannot diff against one.** So Phase 3 was executed against
|
||||||
|
the collapse that *actually exists* — which turns out to be more severe than the premise
|
||||||
|
describes, and measurable right now.
|
||||||
|
|
||||||
|
### 0.1 The collapse points — THREE, not one (VERIFIED)
|
||||||
|
|
||||||
|
**Collapse A — the estimator's components are discarded.**
|
||||||
|
`probabilityEstimator.estimateProbability` returns
|
||||||
|
`{ p_over, p_under, components: { base, recency, weighted, opp_adjustment, home_adjustment,
|
||||||
|
consistency_adjustment, cv } }`. At `analyzeViaEngine1.js:521-524` only the scalar survives:
|
||||||
|
|
||||||
|
const est = estimateProbability({ gameLogs: meta.gameLogs, line: prop.line, ... });
|
||||||
|
const pWin = dir === 'under' ? (1 - est.p_over) : est.p_over;
|
||||||
|
if (pWin != null) legacy.p_win = Math.round(pWin * 1000) / 1000;
|
||||||
|
|
||||||
|
`est.components` is never attached to anything. Discarded: the empirical base rate, the
|
||||||
|
recency rate, each adjustment's magnitude, and `cv` (the only uncertainty proxy produced).
|
||||||
|
|
||||||
|
**Collapse B — THE SEVERE ONE: `p_win` never reaches the grade at all.**
|
||||||
|
The letter comes from `engine1.gradeProp` (`idx = NEUTRAL_INDEX(3); idx += f.delta`), and
|
||||||
|
`engine1.js` contains **zero** references to `p_win` or any probability. So the probability is
|
||||||
|
not "collapsed into the grade" — **it is computed, stored on the payload, and excluded from
|
||||||
|
grading entirely.**
|
||||||
|
|
||||||
|
**Collapse C — `grade_thresholds.json` is read BACKWARDS.** The table maps
|
||||||
|
PROBABILITY → GRADE (`A+ 0.85-1.00`). The live path picks a letter from the factor index and
|
||||||
|
then reads that letter's band MIDPOINT to manufacture `confidence`.
|
||||||
|
|
||||||
|
**NOT produced anywhere today (so not "discarded" — absent):** a posterior distribution, a
|
||||||
|
confidence interval, a similarity prior, an instance count. `cv` is a coefficient of variation
|
||||||
|
of the STAT, not a CI on the probability.
|
||||||
|
|
||||||
|
### 0.2 Market-efficiency-per-sport — SPECCED-BUT-ABSENT, not flattened (VERIFIED)
|
||||||
|
It is not computed-then-flattened; it is **never computed**. One global 11-band scale with no
|
||||||
|
sport dimension. The spec's MLB 0.55 / NBA-stars 0.80 exists nowhere in code. So it is a
|
||||||
|
second *gap*, not a second collapse.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## PHASE 1 — INVENTORY
|
||||||
|
|
||||||
|
| Component | Produced today? | Reaches the grade? |
|
||||||
|
|---|---|---|
|
||||||
|
| Layer-1 similarity prior | **NO** (engine not deployed) | — |
|
||||||
|
| Instance count / min-instance fallback | **NO** | — |
|
||||||
|
| Layer-2 Bayesian posterior | **NO** | — |
|
||||||
|
| Posterior confidence interval | **NO** | — |
|
||||||
|
| Archetype context | YES (snapshot classify) | **NO** — display + challenger only |
|
||||||
|
| `p_over` point estimate | YES | **NO** (Collapse B) |
|
||||||
|
| Estimator components / `cv` | YES, then dropped | **NO** (Collapse A) |
|
||||||
|
| Factor deltas (l5/l20, opp_rank, rest, usage) | YES | **YES — these ARE the grade** |
|
||||||
|
|
||||||
|
**Min-instance fallback (spec: 15): CANNOT DETERMINE / does not fire.** The similarity engine
|
||||||
|
never runs. The code's only abstention rule is `ABSTENTION_RULES.similar_games_below = 3`
|
||||||
|
(`bayesian.py:46`); the only `15` in the file is NBA `min_minutes_per_game`. When data is thin
|
||||||
|
today the live path does not fall back to a prior — `analyzeViaEngine1` **REFUSES** the read
|
||||||
|
(`insufficient_data: true`, grade null) if `projectionFor` finds no reference.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## PHASE 2 — THE DISTRIBUTION → GRADE MAPPING (design, on paper)
|
||||||
|
|
||||||
|
Five explicit rules. **Each is stated so it can be falsified on the forward instrument. None
|
||||||
|
of them is assumed to be an improvement.**
|
||||||
|
|
||||||
|
**R1 — BASE: the posterior sets the letter.**
|
||||||
|
`grade = band(p_posterior)` using `grade_thresholds.json` read FORWARD (the direction it was
|
||||||
|
written for). *Testable:* within each band, realized hit rate should fall inside the band.
|
||||||
|
|
||||||
|
**R2 — UNCERTAINTY GRADES DOWN (the explicit rule, not smuggled).**
|
||||||
|
Compute the posterior's SE; shrink toward 0.5 by the CI half-width before banding:
|
||||||
|
|
||||||
|
halfWidth = z * SE(p) // z = 1.96
|
||||||
|
p_adj = 0.5 + (p - 0.5) * (1 - k * min(1, halfWidth / W0))
|
||||||
|
grade = band(p_adj) // defaults k = 1.0, W0 = 0.15
|
||||||
|
|
||||||
|
So two props at the same `P(>=line)` grade DIFFERENTLY when their uncertainty differs — the
|
||||||
|
wider one grades lower. *Testable, and falsifiable:* rows that R2 moves DOWN should hit closer
|
||||||
|
to their NEW band than their OLD band. If they hit closer to the old band, **R2 is wrong and
|
||||||
|
must be dropped** — high uncertainty would then be noise, not a reason to downgrade.
|
||||||
|
|
||||||
|
**R3 — ARCHETYPE ADJUSTS THE PROJECTION, NEVER THE LETTER.**
|
||||||
|
Archetype modifies the projection *before* the probability is computed; it must never be a
|
||||||
|
post-hoc letter bump. *Testable:* archetype-adjusted projections should lower MAE against
|
||||||
|
actuals vs unadjusted. If MAE does not improve, R3 is dropped.
|
||||||
|
|
||||||
|
**R4 — MARKET-EFFICIENCY SCALES THE THRESHOLD, NOT THE PROBABILITY.**
|
||||||
|
|
||||||
|
threshold_sport(letter) = 0.5 + (threshold_base(letter) - 0.5) * E_sport
|
||||||
|
|
||||||
|
Higher `E_sport` (more efficient market) demands more probability for the same letter.
|
||||||
|
*Testable, and this is the definition of the goal:* after scaling, **the realized hit rate for
|
||||||
|
a given letter should be EQUAL across sports.** That is what "an A means the same thing
|
||||||
|
everywhere" cashes out to, and it is measurable. `E_sport` is FIT from each sport's own
|
||||||
|
accrued record — never hand-set.
|
||||||
|
|
||||||
|
**R5 — ABSTENTION, not a default C.** Below the min-instance threshold, refuse (grade null),
|
||||||
|
matching today's honest refusal behaviour. The threshold is a declared parameter (code says 3;
|
||||||
|
the spec's 15 is unreconciled) and must be set explicitly, not inherited by accident.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## PHASE 3 — THE COLLAPSE COST, MEASURED
|
||||||
|
|
||||||
|
**Method.** 354 SETTLED public ledger rows carrying both the served champion letter and the
|
||||||
|
stored `p_win`. `p_win` was computed pre-game and locked, so this is a forward test, **not
|
||||||
|
lookahead**. The champion letter is compared against the probability letter obtained by
|
||||||
|
reading `grade_thresholds.json` FORWARD (R1 only — R2/R3/R4 are not measurable from stored
|
||||||
|
data, see the gap note below).
|
||||||
|
|
||||||
|
### Discrimination — grade vs realized outcome (point-biserial r)
|
||||||
|
|
||||||
|
| population | n | **champion letter** | **probability letter** | raw `p_win` |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| **ALL** | 354 | **0.0050** (t≈0.09, p≈0.93 — null) | **0.1313** (t≈2.49, **p≈0.013**) | 0.1349 |
|
||||||
|
| **MLB** | 224 | 0.0686 (n.s.) | **0.2356** (t≈3.61, **p≈0.0004**) | — |
|
||||||
|
| **WNBA** | 130 | **−0.0986** | **−0.1258** | — |
|
||||||
|
|
||||||
|
### Hit rate by letter
|
||||||
|
|
||||||
|
| champion | n | hit% | | probability | n | hit% |
|
||||||
|
|---|---|---|---|---|---|---|
|
||||||
|
| A | 2 | 50.0 | | A+ | 24 | 62.5 |
|
||||||
|
| B | 168 | **52.4** | | A | 29 | 65.5 |
|
||||||
|
| C | 174 | **56.9** | | A- | 25 | 44.0 |
|
||||||
|
| D | 5 | 20.0 | | B+ | 30 | **70.0** |
|
||||||
|
| F | 5 | 20.0 | | B / B- | 54 / 48 | 53.7 / 56.3 |
|
||||||
|
| | | | | C+ / C / C- | 43 / 50 / 15 | 53.5 / 50.0 / 53.3 |
|
||||||
|
| | | | | D / F | 26 / 10 | **30.8** / 40.0 |
|
||||||
|
|
||||||
|
**Flat-stake ROI:** probability-top (A-family, n=78) **−1.42%** vs probability-bottom
|
||||||
|
(C-/D/F, n=51) **−26.62%** — a **25-point spread**. The champion's top tier is n=2 (unusable).
|
||||||
|
Letters actually used: champion **3-4**; probability **10-11** of 11.
|
||||||
|
|
||||||
|
### 3.6 — VERDICT: **COSTLY on MLB, and un-collapsing does NOT help WNBA**
|
||||||
|
|
||||||
|
- **The champion's served letter is uninformative.** r = 0.005 overall, and it is **INVERTED
|
||||||
|
between its only two populated tiers — B hits 52.4% while C hits 56.9%.** A user reading B
|
||||||
|
as better than C is reading noise.
|
||||||
|
- **MLB: the collapse is costly and the cost is significant.** The discarded probability
|
||||||
|
carries real signal (r = 0.236, p ≈ 0.0004) where the served letter carries effectively none
|
||||||
|
(0.069, n.s.). **This is the strongest single finding of the model line to date.**
|
||||||
|
- **WNBA: un-collapsing makes it WORSE, not better** (prob r = **−0.126**, i.e. inverse). The
|
||||||
|
pooled r = 0.131 is MLB's signal diluted by WNBA's inversion. This independently corroborates
|
||||||
|
the 2026-07-26 calibration diagnosis ("champion does not discriminate on WNBA"; Brier worse
|
||||||
|
than always-0.5).
|
||||||
|
- **Therefore the full-output challenger must be MLB-FIRST.** Shipping it for WNBA on the
|
||||||
|
strength of a pooled number would ship an anti-predictive grade. WNBA needs a different fix.
|
||||||
|
- Banding costs almost nothing vs the raw probability (0.1313 vs 0.1349) — the 11-band scale is
|
||||||
|
not the lossy part.
|
||||||
|
|
||||||
|
### GAP — R2/R3/R4 are NOT measured here (CANNOT DETERMINE)
|
||||||
|
`SE(p)` needs the per-row instance count `n`, which **is not stored on the ledger**. So the
|
||||||
|
CI-width rule, archetype-on-projection, and efficiency scaling are **specified but unmeasured**.
|
||||||
|
**The challenger build must emit and persist `n`, `SE`, and the pre-adjustment `p`** or R2/R4
|
||||||
|
can never be adjudicated. That is a hard requirement on the next order.
|
||||||
|
|
||||||
|
### 3.7 — Consumers that assume the grade is a POINT
|
||||||
|
Nothing consumes a distribution, so a distribution-based grade is safe **provided it still
|
||||||
|
emits a letter**. What DOES change is the letter DISTRIBUTION, and these depend on it:
|
||||||
|
`tierGating` (Desk gates), `AccuracyBadge` / `ModelRecord` / `TierRecord` byGrade buckets,
|
||||||
|
`outcomeService.gradeBucket` + `TIERS`, `heroPropService.isAB` and `deskShowcaseService.isAB`
|
||||||
|
(A/B filters — the hero pool would grow), `gradeRanking.GRADE_RANK`, `selectTopGrades`,
|
||||||
|
`grade_thresholds.json:capper_minimum_grade = 'A-'`, and the newsletter/media templates.
|
||||||
|
|
||||||
|
**🔴 The A-grade question, stated precisely.** Probability grading would emit **53 A-family
|
||||||
|
rows where the champion emitted 2**. CLAUDE.md's permanent founder ruling forbids *rescaling
|
||||||
|
thresholds to mint A's*. **This is not that** — it is grading on a different and measurably more
|
||||||
|
informative basis (r 0.236 vs 0.069 on MLB), which is new information, and the A's earn it
|
||||||
|
(prob-A 65.5%, A+ 62.5%, B+ 70.0% vs D 30.8%). **But one honest caveat: A- hits 44.0%, breaking
|
||||||
|
monotonicity at the top.** The A-RATED marketing hold should stay until the challenger's own
|
||||||
|
forward record shows a monotone top tier — the ruling's spirit (never sell a relabelled B as an
|
||||||
|
A) is satisfied by evidence, not by the mapping alone.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## PHASE 4 — RE-ADJUDICATION LIST (collapsed-output verdicts, flagged not re-run)
|
||||||
|
|
||||||
|
Every one of these was measured on the collapsed output and must be re-measured on the full
|
||||||
|
output before being trusted forward:
|
||||||
|
|
||||||
|
1. **The champion's p_win→CLV edge** (partial r = 0.375, p ≈ 0.003, n = 62 takeable MLB overs)
|
||||||
|
— measured on the collapsed p_win. Direction is likely preserved (it already used p_win, not
|
||||||
|
the letter), but the magnitude is a collapsed-output number.
|
||||||
|
2. **The over-side skew audit** ("SURVIVES BASELINE", +7.14pt marginal) — same basis.
|
||||||
|
3. **proj-v1.1's "NOT PROVEN"** — it was judged against the COLLAPSED champion. A full-output
|
||||||
|
champion is a different, stronger benchmark, so **proj-v1.1's death is not final** — it must
|
||||||
|
be re-run against the new champion, and could fare better or worse.
|
||||||
|
4. **The takeable-floor derivation (C1)** — ROI-by-price buckets were computed on rows graded by
|
||||||
|
the collapsed engine; the population selection itself is collapsed-output-conditioned.
|
||||||
|
5. **The calibration diagnosis (2026-07-26)** — Brier/reliability/resolution on collapsed p_win.
|
||||||
|
Its WNBA finding is *corroborated* by this order and looks robust; its MLB numbers need re-running.
|
||||||
|
6. **ROI-by-grade (MLB-C +4.57%, MLB-B −1.02%, WNBA −5%)** — these are grade buckets from a
|
||||||
|
letter now measured at r ≈ 0.005. **Given B and C are inverted, "MLB-C is the profitable
|
||||||
|
segment" is very likely an artifact of a meaningless letter, not a real segment.**
|
||||||
|
7. **`confidence` on every historical row** — a band-midpoint of a letter with no predictive
|
||||||
|
content; it should not be used as a weight in any future analysis.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## TAGS
|
||||||
|
VERIFIED: the three collapse points; the engine is not wired/deployed; discrimination numbers
|
||||||
|
(n=354, per sport); consumer list. CANNOT DETERMINE: R2/R3/R4 effects (per-row `n` not stored);
|
||||||
|
the spec's min-15 instance rule (absent from code). BLOCKED: none.
|
||||||
|
|
||||||
|
## HELD
|
||||||
|
No grade change, no challenger build, no promotion, no reconnection, no retro-grading.
|
||||||
Reference in New Issue
Block a user