Commit Graph

7 Commits

Author SHA1 Message Date
builtbykev f7cc19772b Canonical MLB participant: prove the human, then dedupe
VYNDR had three definitions of "same player": raw p.player (dedupe),
utils/normalize.normalizeName (features, NO nicknames), utils/playerName.nameKey
(retention, WITH nicknames). Retention was the first layer to notice they
disagreed, and all it could do was discard the loser — 28 collisions.

The fix is not a better spelling. The event-scoped roster ALREADY carries the
StatsAPI personId; buildPlayerTeamIndex was fetching and discarding it.

SCOPE IS LOAD-BEARING, and this is the finding that shaped the design. Measured
on the real 2026-08-28 league rosters (1,583 rows, 30 teams): a bare name key is
NOT globally unique — `max muncy` (ATH 691777 / LAD 571970), `jose fermin`
(665877/820862) and `luis garcia` (472610/671277) each resolve to TWO different
humans. Team-scoped: 0 ambiguous. Event-scoped across 33 real events including
the verified doubleheader date: 0 ambiguous. So resolution is scoped to the two
teams actually playing, and FAILS CLOSED — no event, no game, no candidate, or
more than one candidate yields null and the prop keeps prior behaviour.

FEATURE IDENTITY PARITY, measured on both real cohorts before writing the patch
(504 participants):
  RAW_SUCCESS/CANONICAL_SUCCESS/SAME id      365
  RAW_SUCCESS/CANONICAL_SUCCESS/DIFFERENT id   0   <- required 0
  RAW_SUCCESS/CANONICAL_FAIL                   0   <- required 0
  CANONICAL_AMBIGUOUS                          0   <- required 0
  RAW_FAIL/CANONICAL_SUCCESS                   7   <- repair, not regression
The seven improvements are exactly the alias class: Michael->Mickey Gasper,
AJ->A.J. Ewing, JT->J.T. Realmuto, Mike->Michael Busch, Richard->Richie
Palacios. This is why raw spelling could not be trusted: player_id_map holds
`mickey gasper` and NOT `michael gasper`, so under the raw name the lookup
outcome depended on which alias happened to survive dedupe. Provider arrival
order was deciding model input availability.

ALL SEVEN COLLISION GROUPS PROVEN SAME_PLAYER_ALIAS against the authoritative
roster for the exact event date — one personId each (681715, 699625, 676356,
695491, 673357, 681508, 671739), zero false normalizations, zero unresolved —
and all seven collapse to one participant under the new key, so the 14 colliding
propositions merge upstream instead of being discarded downstream.

FALSE-MERGE SIMULATION on both cohorts: cross-event 0, different-stat 0,
different-line 0, false merges 0. Note honestly: the retained rows cannot
exhibit the merges themselves, because retention already discarded the losers —
so the merge half is proven directly on the alias groups, the safety half on the
cohorts.

RETENTION IS UNTOUCHED — schema, player_key, conflict identity and the collision
counter are all unchanged. collision_count must reach 0 by upstream repair, never
by making the counter lenient. A teeth proof injects raw spelling into the
retention identity and the suite rejects it.

Source provenance preserved: p.player is never overwritten; the participant
rides beside it as mlb_person_id + canonical_player_name.

Twelve teeth against a green baseline of 117 — dedupe back to raw name (3),
feature lookup back to the alias (2), resolver stops blocking ambiguity (1),
lookup degraded to raw (1), event scope dropped (1), event/line/stat dropped
from the key (2/2/2), provenance destroyed (1), raw spelling in retention
identity (2), collision_count accepted nonzero (2), and a guard proving the
frozen game-date repair cannot be reverted (10). Restored byte-identically.
Teeth #4 (canonical resolving to a DIFFERENT feature id) is covered by the
real-data parity measurement rather than an injected branch, because it is a
data comparison, not a code path — stated plainly rather than claimed as a test.

No hardcoded player names in production code — a test greps for all seven.

4 files, 86 insertions. retentionService, gameBinder, oddsService,
acquisitionTrace, probabilityEstimator, bookRoles, playerName, normalize and
snapshotScheduler: UNCHANGED. Admission rules, MODEL_BOOK filters and model
formulas: 0 changed lines. The game-date repair is intact.

391 suites / 5,339 tests pass. web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 12:54:35 -04:00
builtbykev 6415751f2e Grading binds opponent features to the REAL game, not ESPN's "today"
Order 1.6 Phase 1. This is a MODEL-OUTPUT fix, not bookkeeping.

computeFeatures.lookupTodayGame called the ESPN scoreboard with NO date
param and took whatever ESPN calls "today". Renamed to lookupGameOnDate
and now sends ?dates=YYYYMMDD from the prop's BOUND game — the same game
the ledger, retention and settlement use, so all four finally agree.

PROVEN against live ESPN (before/after, same instant):
  dateless "today"      CLE->PIT  NYY->LAD  LAD->NYY   (Jul 19 card)
  bound to 2026-07-20   CLE->MIN  NYY->PIT  LAD->PHI   (the real games)
  bound to 2026-07-19   CLE->PIT  NYY->LAD  LAD->NYY   (reproduces OLD)
Every opponent was wrong. opponentAbbr feeds opp_rank_stat (a +/-1.0
factor) and isHome feeds home_away (+0.5), so late-slot grades were
scored against the wrong matchup.

Note the window is WIDER than the 01:00/03:00 UTC slots: this ran at
07:5x UTC = 03:5x ET and ESPN's dateless scoreboard was STILL returning
the previous day's card.

HONEST DEGRADATION: with no bound game date the grader does NOT fall back
to a dateless lookup — it records 'no_bound_game_date' and leaves
opponentAbbr/isHome/gameId null, so engine1 simply omits the opponent and
home/away factors rather than scoring a wrong matchup. Tests assert both
directions.

Same class of bug fixed alongside: the Tank01 augmentation used TODAY's
UTC date for its cache key; it now uses the bound game date.

gradeSlateService threads game_date/game_time/home_team/away_team into the
grader so the binding reaches computeFeatures at all.

Audited the rest of the feature path for dateless/"today" lookups — none
remain (weather is current-conditions by venue, park/pace are static).

Suite 282/3383 green, build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-20 03:50:07 -04:00
builtbykev 1a94ef5fcf Revive the dead probability layer + restore grade range ON MERIT
Folds re-sequenced steps 1+2 into one change (Kev's call): same bug
family — features wired to sources that return null.

THE PROBABILITY LAYER WAS DEAD IN PRODUCTION. p_win/ev_pct/kelly/
model_odds/value were absent on 0/8 live grades because
gameLogService.getGameLogs returns null for MLB by construction and
depends on the offline Python service for NBA/WNBA, so meta.gameLogs was
[] for every sport. This was the S46 bug in a second location — that fix
gave featureCache an MLB branch (why grades still worked) but never the
estimator. featureCache.getStatRows now supplies normalized rows
([{date,[statType]:v}], most-recent-first) for every sport, feeding the
estimator AND consistency AND game_count_in_7d from one fetch.
VERIFIED on real props: p_win 25/25 WNBA, 8/8 MLB (was 0).

GRADE RANGE, ON MERIT — never by rescaling (permanent founder ruling:
minting A's without new information is a relabelled B sold as an A and
corrupts an append-only ledger).
- refreshTeamStats wired into runSnapshot — it had ZERO production
  callers, so opp_rank_stat was permanently null and a +/-1.0 factor
  could never fire. Test-env no-op (opsNotify precedent).
- L20 made SYMMETRIC: both branches were delta +1.0, so the season
  baseline could only ever ADD. No negative path was a structural reason
  D was unreachable. New l20_contradicts_* carries -1.0.
- game_count_in_7d derived from real logged dates (heavy_workload_7d).
- NOT wired, deliberately, with reasons inline: teamId (no team_id
  column; getFeatures reads it top-level; factor also needs a starter-id
  list) and season_type (ESPN 2 = REGULAR season; threading it raw would
  fire veteran_in_playoffs in July). Dead code dressed as a fix is the
  thing we are removing, not adding.

CALIBRATION GUARD (found by verifying, not assuming): consistency CV is
NBA-tuned; for a Poisson-ish stat cv ~ 1/sqrt(mean), so any stat with
mean < 4 auto-classifies boom_bust. First verification run showed 8/8 MLB
props boom_bust — a blanket -1.0 that dropped the board to all-C. Floored
at CONSISTENCY_MIN_MEAN=4 -> 'unknown' below. Absent beats wrong. MLB
low-count stats therefore still get no consistency factor: honest, not
fixed. Scale-free index-of-dispersion classifier is the open follow-up.

CONFIDENCE IS NOT A PROBABILITY: payloads carry confidence_basis:
'grade_band'. Corrected mlb-grade-degradation.md — its "25/25
grade<->confidence agreement" is a TAUTOLOGY (confidence is derived FROM
the letter, so it would report 25/25 even if every grade were wrong), not
a validation. Removed dead mlbGrader.js (referenced only by its own test)
and the stale computeFeatures comment claiming a penalty that never ran.

VERIFICATION (scripts/verify-grade-range.js, real props/logs/engine):
WNBA 25 props B 68%->32%, C 32%->64%, D 0->1 (4%); 11-step spread went
from 2 steps to 5 (C/C+/B-/D). The D is earned: Angel Reese assists o2.5,
p_win 0.365. Nothing flooded — grades got HARDER. A did not emit locally
because opp_rank_stat needs the Redis cache only prod populates (local
ceiling +3.0 vs the +4.5 A needs); reachability is proven arithmetically
and locked in tests. Prod A-emission is the outstanding fingerprint.

MARKETING HOLD: "A-RATED" (AccuracyBadge, TopSignals) is unsupported
until that fingerprint. Confirmed honest fallbacks render today —
/api/ledger/accuracy has B and C buckets only, so the badge shows
"MODEL · 63% HIT" and TopSignals self-hides. Nothing fabricated ships.

Suite 276/3286 green, web build exit 0.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
2026-07-19 18:54:51 -04:00
builtbykev 167996d99a Session 15: Intelligence hardening — park factors, weather, Tank01 prefetch, pace factors, signal audit, founder pricing fix (1405 tests)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-11 16:21:18 -04:00
builtbykev f5d79cf70d Session 14: Africa checkout, Tank01 NBA/MLB wiring, WNBA+MLB odds proxies, OAuth icons, loading skeletons (1330 tests) 2026-06-11 10:06:49 -04:00
builtbykev ad5ea8d5a8 Session 7j: Soccer intelligence - 9 leagues, 11 signals, 6 traps, poller, prefetch, 131 new tests (1173 total) 2026-06-10 14:50:13 -04:00
builtbykev 4815ceac03 Sessions 7e+7f: Grade adapter, normalize consolidation, computeFeatures, analyzeViaEngine1, scan/parlay migrated to engine1 2026-06-10 09:28:30 -04:00