Commit Graph

319 Commits

Author SHA1 Message Date
builtbykev 84f1fc075c The live machinery ships dark, behind two gates that cannot substitute for each other
ATTRIBUTION RECEIPT (cohort 27ce152f, writer 94f7c3c, 01:02:19Z): 2,316 non-hits
rows — certified 0, numeric 0, and artifact_id / estimator / certification
version / source / knot / curve ALL ZERO. The hits slice held: 209 published,
201 certified, curve mismatches 0/201, artifact identity deviations 0, support
violations 0, raw erased 0. Shadow tranche frozen.

AUTHENTICATED NON-LEAK: obtained through the owner magic-link flow — the real
auth system, no bypass, no stored password, token never printed, session
discarded. /api/ledger, /api/ledger/accuracy, /api/preferences, /api/accuracy
and the authenticated /api/snapshot/mlb (677KB, full model fields) carry zero
shadow keys. One scanner hit adjudicated: `served_grade.calibrated` is false on
all 504 rows — servedGrade's hardcoded constant from before this work, a name
collision with my keyword list, not the shadow.

A REAL MONITOR DEFECT, asked for and found. The row floor alone let 400
observations from ONE slate reach HEALTHY or DRIFT — one correlated draw wearing
the costume of four hundred. The monitor now also requires settled DATES, and
the floor is not invented: it reads
`fitPolicy.POLICY_V1.certification.eval_block_dates`, the 3-date fold the
walk-forward was actually certified with. One date and two dates now return
INSUFFICIENT_SAMPLE regardless of row count.

That fix had a bug of its own that a test caught: `scored` never carried `date`,
so the distinct-date count read `undefined` for every row and always returned 1.
The gate looked correct while measuring nothing.

THE SEAM. EV, Kelly and VALUE are produced in exactly one place at grade time,
all from p_win, and every served row passes exactly one boundary on the way out
(`snapshotGating.stripModelPrice`, used by the snapshot route, hero route,
topGraded and props). So `servedProbability.applyToRows` sits there — one place,
not four call sites and four chances to miss one. Consumers never see the
artifact, the curve, the support region or the environment; a test forbids
`applyCurve`, `applyIsotonic`, `artifactRegistry` and `served_curve` in all three
serving consumers.

Placed BEFORE the tier strip deliberately: calibration decides what the number
IS, entitlement decides who may see it. Reversed, it would calibrate fields that
had already been removed.

TWO GATES, NEITHER SUFFICIENT. The artifact is promoted APPROVED_FOR_LIVE — stage
only, same id, same source/knot/curve digests, same cutoff, no refit. Behaviour
is still OFF because PROBABILITY_CONTRACT_LIVE is unset. Flag without approval:
OFF. Approval without flag: OFF. Both: ON. Live OFF returns rows byte-identical
and consults no artifact.

Two teeth had to be rewritten rather than satisfied: they pinned "artifact
unapproved" as the safety property, which would have blocked the deliberate
promotion. The property is that live BEHAVIOUR is off, and they now test that.
And my own splice while fixing one of them silently deleted ten newly-added
teeth — caught by counting ids, not by the runner going green.

Suite 406/406, 5,681 passed. Teeth 41/41 + 10/10 + 23/23. Live OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-04 06:36:11 -04:00
builtbykev 94f7c3c3ef The probability never leaked across stats; the attribution did
Post-fix cohort df4ec562 closed the primary question: eleven non-hits stats,
1,910 rows, ZERO certified and ZERO numeric served probabilities. The cross-stat
repair holds.

But every one of those 1,910 rows still recorded
`artifact_id: mlb-hits-isotonic@2026-09-03` beside state UNSUPPORTED. A `doubles`
row named the hits artifact. Nothing was calibrated by it, so no number leaked —
but a later query for "rows this artifact produced" would have returned 2,186
instead of 133, and that is the shape of footgun this programme keeps finding.

The contract check runs BEFORE any artifact is relevant: with no certified
contract for the sport/stat, no artifact applies, and naming one asserts a
relationship that does not exist. UNSUPPORTED now carries null artifact,
artifact_id, estimator_type, estimator_version, certification_version and
procedure_version.

Attribution is KEPT where the artifact is genuinely the thing that declined —
UNCERTIFIED (out of support) and VERSION_MISMATCH both still name it. A test
holds both directions so this does not over-correct into erasing real provenance.

One existing test called resolve() without naming a stat and relied on the
service substituting one. That substitution was the original defect, so the test
now names its stat, as production does.

Artifact unchanged. Live OFF. Suite 405/405, 5,662 passed. Teeth 35/35 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 18:33:31 -04:00
builtbykev 981b26010a The first real cohort found the leak: one artifact was calibrating every stat
Shadow converged at 18:09:07Z and the 19:00Z slot produced cohort 0353c551.
It immediately falsified something no test had asked: 513 PUBLISHED NON-HITS
rows came back CERTIFIED_CALIBRATED with a served probability drawn from the
mlb-hits curve — total_bases 246, runs 149, rbi 125, walks 109, outs 28,
strikeouts 24, hits_allowed 20, earned_runs 17.

Two bugs, one on top of the other. `mergeProbabilityContract` passed only
{model_version, p_win}, dropping the row's identity; and the service's resolve
then stamped the {sport, stat} it had been BUILT with onto every read. So all
3,000 rows in the batch resolved as mlb hits.

The governance tests could not see it. They asked "does build() refuse another
stat?" — it does, and always did — and then exercised the merge with
hits-only rows. Production sends one mixed batch. The regression test now drives
the REAL collector with hits, total_bases, rbi, runs, walks, strikeouts and
home_runs at the same p_win and requires hits certified and every other stat
neither certified nor numeric.

Fixed in three layers, because one would have been the same single point that
just failed:
  1. the service no longer substitutes its own identity — the row's decides,
     and a read naming no stat resolves to no contract, which is UNSUPPORTED;
  2. the merge carries the row's sport and stat;
  3. probabilityContract refuses an artifact whose own sport/stat disagree with
     the contract it is being used under, independent of plumbing.

NO USER IMPACT. Shadow only: every block carries servable:false, live serving is
OFF, CALIBRATION_DEPLOYED is [], and the anonymous payload showed zero
calibration fields before and after. But this is exactly the defect that would
have served a hits calibration curve for strikeouts on the day live was enabled,
and only a real cohort surfaced it.

Two teeth were themselves wrong. Both runners checked "retention identity
changes" by grepping the diff for `stat:`, which fired on `stat: r.stat` — a
line that READS identity to hand it to a reader, not one that changes what
identifies a row. A guard that cannot tell those apart blocks the fix for the
defect it exists to protect against. Both are now behavioural: build a row
through the real collector and compare the identity tuple.

Artifact unchanged: mlb-hits-isotonic@2026-09-03, knot 5ae940ea163b7da2.
Suite 405/405, 5,659 passed. Teeth 34/34 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 16:07:39 -04:00
builtbykev 044df7a426 The monitoring contract gets a callsite
Traced 2026-09-03: `calibrationRegistry.reverify` had ZERO production callers,
and the only reference to the artifact machinery outside its own directory was
the shadow builder, behind a flag that is off. The previous tranche declared
that future settled outcomes are forward evaluation evidence and shipped no
evaluator — the contract lived in a comment. That is the same shape as
statModel.js and correlateValidator.js, cited for months and never present.

`forwardMonitor` scores the FROZEN artifact on outcomes settled strictly after
its training_cutoff, and `snapshotScheduler.calibrationMonitorTick` runs it on
the existing per-minute cadence, throttled 6h, every failure swallowed. A test
asserts the tick is defined, invoked AND exported, and drives it with a fake to
prove it passes the promoted artifact and a forward-only window.

It cannot change what it watches: the module imports no fitter and no registry,
and a test greps the stripped source for fitIsotonic, fitPlatt, writeFileSync,
upsert, update, promote( , artifactRegistry and PROMOTED.

LOW N IS ITS OWN ANSWER. Below the floor it reports INSUFFICIENT_SAMPLE with
`healthy: null` — never false, never true. The floor is DERIVED, not chosen:
resolving an error of the certified tolerance at two standard errors needs
n >= 0.25/(0.05/2)^2 = 400, and a test recomputes it from TOLERANCE so the two
cannot drift apart. TOLERANCE is 0.05, the same number certifyBands used, so the
monitor can be neither stricter nor laxer than the thing it watches.

Wrong-era forward rows are INVALID, not scored — scoring them would measure a
different forecaster, which is the defect this line of work removed. An
unreadable read is AUDIT_UNAVAILABLE, which is not a health verdict. A drift
alert says in its own text that the artifact is frozen and unchanged, because
the alert is not a demotion.

`currentEraSource.loadRows` gains an optional `after` bound, strictly greater
so the cutoff date itself can never be scored as forward evidence.

SHADOW REMAINS BLOCKED: the production variable is still absent
(configuration_source "default") after a restart at 05:09:02Z, so no shadow
cohort was obtained and no replay was substituted for one.

Artifact unchanged: mlb-hits-isotonic@2026-09-03, knots 5ae940ea163b7da2.
Live OFF. Suite 405/405, 5,654 passed. Teeth 31/31 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:29:47 -04:00
builtbykev f4f51ed7ca The runtime names its artifact before anything is switched on
Step 35 asks the deployed process to prove WHICH frozen curve it resolves —
artifact id, procedure, era, source digest, knot digest, curve digest, support,
servable, stage — with the shadow still OFF. The probe could not answer, so the
same shadowState() evaluator now resolves the promoted artifact and reports its
identity. Identity only: the curve is committed and addressable by artifact_id,
and repeating it on a status page would put a second copy of the truth there.

null means nothing is promoted, which is also the honest answer when the file is
missing or fails its policy at load.

Suite 404/404. Shadow OFF. Live OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:08:44 -04:00
builtbykev be8e16aca9 Detection becomes repair: the curve is fitted on one forecaster, frozen, and named
The last release detected the violation and then served the certified state
anyway. A validator that changes nothing is decoration, so `servable:false` is
now load-bearing: an artifact that fails its policy returns
ARTIFACT_POLICY_BLOCKED with no number, and every probability-derived claim goes
with it. The gate sits inside the resolution, not beside the flag that turns the
shadow on, so no environment variable can reach past it — a test asserts
`resolve` never reads process.env at all. Shadow and live consume the SAME
decision, differing only in which promotion stage they demand.

Era mismatch still resolves to VERSION_MISMATCH rather than the new state. "This
artifact belongs to a different forecaster" is more precise than "policy
blocked", and the existing state already says it exactly.

THE REPAIR. `currentEraSource` filters on model_version in the QUERY, taking the
era from config/modelVersion so the query, the artifact and the validator all
read one identity. Measured on the actual fitted set, not a second count:
6,069 current-era rows, 0 wrong-era.

The procedure was then certified on current-era rows ONLY — four walk-forward
folds, training strictly before each evaluation block, 0 future rows in train on
every fold. All four improve; pooled n=3,108 gives Brier 0.24701 -> 0.24323,
delta -0.00378, CI [-0.00619,-0.00147] excluding zero; ECE falls in every fold.
Mapping spread inside support is 0.001-0.018. The prior mixed-era certification
did not substitute for this.

Policy B selected. A (era-filtered 65/35) and B (all current-era) are
statistically indistinguishable, A-B = +0.0001 CI [-0.00029,+0.00048], but B has
the better ECE (0.0064 vs 0.0109) and the holdout existed to certify the
PROCEDURE — it is not permanently withheld from the artifact that ships.
withheld_from_fit is 0.

FROZEN. `mlb-hits-isotonic@2026-09-03`: 6,069 rows, training_cutoff 2026-09-01
(distinct from fit_as_of 2026-09-03 — the newest observation admitted is not the
eligibility bound), 12 knots, source_digest 25919c16…, knot_digest 5ae940ea…,
served_curve_digest c24a9dc5…, 8 curve steps, 924 bytes, committed as JSON.

The runtime no longer fits. It loads. A test greps the service for fitIsotonic,
fromLedger and loadRows and requires all three absent, because the old behaviour
meant a user's number could move with no version, no review and no rollback, and
a past Read could not be reconstructed because its curve no longer existed.
New settled outcomes are forward evidence now; they cannot touch this curve.

Independent reconstruction from the declared training contract alone — fresh
read, fresh digest, fresh fit — reproduces every digest and the curve byte for
byte. Calling the builder twice would only have proven the builder deterministic.

Promotion is a frozen source constant. A snapshot cannot promote, a settlement
cannot promote, a successful fit cannot promote, and dropping a file into the
artifacts directory promotes nothing. Stage is APPROVED_FOR_SHADOW; live is
explicitly false.

Two coverage holes found by their own teeth. The promotion guard could be
deleted with every test still green, because the promoted file naturally agrees
with itself — extracted as `acceptFile` and tested on the case `load()` cannot
reach. And `validate(null)` returned no `servable` field at all, which is falsy
at a call site and so would have read as correct while asserting nothing.

Shadow OFF. Live OFF. CALIBRATION_DEPLOYED []. No frontend change.
Suite 404/404, 5,634 passed, 4 skipped. Teeth 26/26 + 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-03 01:06:18 -04:00
builtbykev 5cad851922 A digest names the curve; it does not vouch for the procedure that makes the next one
Artifact identity was the previous tranche's answer. It is not certification.
`knot_digest` says WHICH mapping ran and nothing about whether tomorrow's refit
deserves the same trust — that one would get its own digest and be equally
"identified" while fitted on anything at all.

So the object being certified is named: VYNDR certifies a PROCEDURE, not a
frozen curve. A frozen curve goes stale against a live model and has to be
replaced by hand on no schedule; "fit past, apply forward" is procedural by
construction. `fitPolicy.POLICY_V1` declares it — data selection, horizon,
algorithm and version, minimum rows, model-version restriction, sport, stat, and
a support contract that a refit may NOT widen. Each artifact still carries its
own digest.

`fitPolicy.validate` is the Step-22 gate: an artifact does not become servable
because the algorithm ran. It refuses a widened support, a wrong era, a wrong
estimator, a thin fit, a missing identity or a missing training cutoff, and
`servable` is false whenever any violation stands, with no override argument.
The statistical bars stay where they already live in calibrationRegistry — this
is not a second governance system.

MEASURED, AND THE REASON LIVE SERVING STAYS BLOCKED: production does not match
the declaration. `loadSettledRows` applies no model_version filter, so at
fit_as_of 2026-09-02 the fit drew 6,084 rows from 9,361 settled — all 3,292 from
the superseded engine1@2026-07-20 plus 2,792 current-era. 54.1% of the served
map is fitted on a forecaster it was never certified for, while the artifact
declares the current era.

That is a provenance contradiction, not a performance claim: era-filtered scores
0.24065 against pooled 0.24068 on 1,120 out-of-sample rows and both intervals
span zero. It is blocked because nothing prevents the next era change from
repeating it, and because the freshness lag grows.

`era_restricted` is answered STRUCTURALLY, not by an extra read — the query is
in this service and applies no filter, so the artifact records
ERA_NOT_RESTRICTED rather than claiming a restriction that did not hold. An
unverified restriction is recorded as a violation, because "we did not check" is
exactly the state production is in.

The violation does NOT distort the shadow. Flipping every row to UNCERTIFIED
would make the shadow measure the violation instead of the contract, so the
policy state rides beside the resolution and a test asserts the shadow still
reads CERTIFIED_CALIBRATED at 0.65 and UNCERTIFIED at 0.91.

Nothing serves. CALIBRATION_DEPLOYED still []. Shadow still OFF (the production
variable remains absent — the probe reads configuration_source "default").

Suite 402/402, 5,607 passed, 4 skipped. Teeth 10/10 + 23/23.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 23:43:31 -04:00
builtbykev 556d186ff1 One evaluator for the shadow, so the probe cannot report a mode the pipeline is not in
"Is the shadow effective?" was only answerable by waiting for a snapshot to
write a row. That leaves a blind spot with real cost: a variable SET IN COOLIFY
BUT NOT YET APPLIED to the running process is indistinguishable from an unset
one, and the runtime probe already proves the distinction matters — code_sha
22cf51c with started_at 01:45:44Z means anything set after that is not in this
process's environment.

probabilityContract.shadowState() is now the single evaluator. snapshotService
calls it and the protected status probe calls it, and a test asserts NEITHER
reads process.env directly — the same rule that keeps lineage_write_mode honest.
Reading the env in two places is how a status page and a gate come to disagree.

Strict by construction: only the exact string '1' enables it. 'true', 'yes',
'on', '01', ' 1 ' and '' are all OFF, because a loose parse turns a typo into an
activation. `configuration_source` separates an unset variable from one
explicitly set to '0', and `live_serving` is reported as its own switch so the
shadow can never be read as implying serving.

No behaviour changes. The shadow still defaults OFF, CALIBRATION_DEPLOYED is
still [], and served fields are untouched.

Frontend byte-identical to the last green build (git reports zero changes under
web/), so the build from 22cf51c stands.

Suite 401/401, 5,597 passed, 4 skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 22:53:21 -04:00
builtbykev 22cf51c4b0 The mapping that ran now has a name, and the row carries it
Runtime probes say the fleet is on a8de676, one process generation. But
"isotonic" was a label, not a claim: probabilityContractService refits per
snapshot against `game_date < todayEt()`, so the mapping changes as outcomes
settle, and nothing on a row could say WHICH mapping produced its number.

The artifact now has an identity:

  estimator_type / estimator_version / certification_version / model_version
  fit_as_of         the exact lt(game_date) bound   2026-09-02
  training_cutoff   last date INSIDE the fit        2026-08-21
  fit_n / knot_count                                6,084 / 28
  knot_digest       d9d571d728ba76de
  served_curve      the COMPLETE served function over [0.50,0.80)
  served_curve_digest

The served curve is not a sample. p_win is quantised to three decimals at the
source, so a step table at 0.001 granularity is the mapping itself for every
input that can occur — six steps, ~200 bytes. Storing it makes a Read
reconstructable WITHOUT re-deriving a training set that may since have been
re-settled, and a claim you can only verify when the inputs happen not to have
moved is not a reconstructable claim.

Proven, not asserted: the production construction path run twice gives an
identical digest, and an INDEPENDENT reconstruction — re-walk 9,361 settled
ledger rows at the declared bound, refit from scratch — reproduces
d9d571d728ba76de exactly, 28 knots for 28.

A teeth injection found a real defect behind a coverage hole. `resolve` checked
the CONTRACT's model era and never the ARTIFACT's, so a mapping fitted for a
different era could be recorded beside a served number with every test green.
Both the era and the estimator type are now checked, and a mismatch serves
nothing rather than serving quietly.

OBSERVED AND NOT CHANGED: calibrationService splits 65/35 to certify its own
bands, a step this contract does not consume because support comes from the
frozen artifact. So the served map is fitted through 2026-08-21 while 3,277
more recent settled rows sit unused, and that lag grows with history. Changing
it would change the fitted function, which this tranche froze.

Shadow still defaults OFF. CALIBRATION_DEPLOYED still []. served_probability is
referenced by nothing outside the contract layer — asserted by a tooth.

Suite 401/401, 5,593 passed, 4 skipped. Teeth 23/23 (prior) + 7/7 (new).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:44:52 -04:00
builtbykev a8de676756 A probability is served because evidence supports it, not because nothing else answered
The band gate was blocked for its `else` branch. It read:

    candidate = F(raw)
    served    = inCertifiedBand(candidate) ? candidate : RAW

and above raw 0.60 the model is measured overconfident — holdout raw 0.80-0.90
predicts 0.843 and realizes 0.639. So "the calibrator is not supported here" was
being answered with a number already proven wrong. Unsupported calibration does
not make raw true.

Four candidates were adjudicated on ONE split — fit on the earliest 60% of
train, decide support on the last 40%, evaluate on a holdout that saw neither:

  A low-param      80.2% coverage  0.24374  REFUTED — its extra region
                   (raw 0.80-0.90) certified on cert (err +0.040, n=55) and
                   refuted on holdout (served 0.754 vs observed 0.639), and it
                   leaves a hole at 0.70-0.80 while serving the island above it
  B isotonic       91.3% coverage  0.24337  CERTIFIED, contiguous raw [0.50,0.80)
  C empirical band 91.3% coverage  0.24335  REFUTED — refitted point-in-time on
                   current-model hits the realized rates INVERT in grade order
                   (B+ 0.593 < B 0.614 < C+ 0.623), so the served function steps
                   down at raw 0.78. Its shipped constants come from 3,417 props
                   pooled across four batter stats and do not reproduce here
  D raw identity   43.1% coverage  0.24866  certifies raw 0.50-0.60 and only there

Raw is candidate D, not a fallback. It earns exactly one region (holdout error
+0.010 on n=1,316), which is why the law is "raw must earn its region" rather
than "raw is never true". B already covers that region, so no hybrid is built.

Above raw 0.80 nothing is certified and nothing is served. That is the region
where raw is most wrong, isotonic over-corrects (cert err -0.093) and its LODO
mapping at 0.95 has spread 0.180. 8.7% of holdout rows land there.

The registry did not need changing. `serves(stat, p)` already tested certified
bands against the RAW p_win — support in the input domain, the correct question —
and returned {serve:false, reason}. It never said "serve raw". The output-space
gate and the raw fallback were both invented downstream in calibrationService.

ACTIVATION IS OFF. PROBABILITY_CONTRACT_SHADOW defaults to 0, CALIBRATION_DEPLOYED
stays frozen empty, and every served field is byte-identical. This releases the
support first, which is the required order. The shadow records raw belief, the
candidate served value, the state, the estimator identity, and what EV/Kelly/VALUE
would be under the actionability law — into its own column, read by nothing.

Migration 051 was applied to production BEFORE retentionService named the column.
PostgREST builds a bulk insert from the first row's shape, so a key whose column
does not exist 400s the whole batch silently — that is how migration 038 took
retention down for three days.

The user-facing contradiction is NOT fixed here. A B+ still says "realized about
66%" beside a confidence of 84. Fixing that is activation, and activation costs
32% of VALUE flags and 46% of Kelly recommendations on the holdout.

Suite 401/401, 5,580 passed, 4 skipped, deterministic across three runs.
Teeth 23/23, each independently injected and restored byte-identically.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-09-02 21:03:16 -04:00
builtbykev 9dda9df132 Read History: the graph becomes something a person can read
The ancestry contract existed and nothing rendered it. This turns it into a
product surface, and stops there.

WHERE IT GOES. `LedgerCard` is the terminal surface — there is no Ledger detail
view — so the history expands in place inside the card, matching the board's
existing "ALL N READS" affordance rather than adding a page, a modal or a
navigation category. It loads on first open, not on render.

WHAT IT IS CALLED. "Read History". Lineage stays engineering vocabulary; a test
asserts no rendered string contains lineage, natural key, ordinal, digest,
graph, origin or recapture. ORIGIN reads "First published", REVISION reads
"Updated", and the persisted `change_type` supplies "The price moved" / "The
read changed" / "The read and the price changed". `change_type` is stored and
trustworthy, so naming it is reporting; no field-level diff is persisted, so
none is invented.

THE SEPARATION, WHICH IS THE LOAD-BEARING PART. The grade strike means the
LETTER changed. A history entry means the published CLAIM changed — often the
price, sometimes the read, frequently with no letter change at all. The history
uses no strike-through, shares no styling, and a tooth fails if it ever does.
The ledger card's own strike is untouched.

RECAPTURES ARE SUMMARISED, NEVER DESTROYED. A republishing board can produce
hundreds of "unchanged" entries that bury the two that matter, so the UI
collapses them to a count. The API still returns every one.

WHAT EACH ENTRY SHOWS came from the acceptance run: without the published grade
the history can say a Read changed but never what it changed to, which answers
none of the questions someone opens a history to ask. `published_grade` /
`published_p_win` / `published_line` are read off the SAME retained row the
lineage action sits on — the authoritative record of the published claim, with
lineage only the pointer to it. A tooth fails if they are ever synthesised from
lineage metadata.

TWO DEFECTS THE PRODUCTION ACCEPTANCE FOUND, NEITHER OF WHICH A TEST HAD.

`ledger_entries.id` is a UUID and the route parsed it with Number.parseInt, so
every real row would have 400'd. The unauthenticated probe that "proved the
route was live" returns 401 from requireAuth before the handler runs, so it
could never have seen this.

And `chase burns / hits_allowed / under / 4.5` has SIX published captures and
four lineage actions — two were published while the writer was off. The response
said `chronology_complete: true` while showing four of six. Completeness now
counts published-but-never-recorded states as well as attempted-and-failed ones,
kept as separate numbers because the causes differ and the copy says which.

Suite 399/5,544/0 · tsc 0 · web build 0 · teeth 15/15.
Lint is not runnable in this repository (`next lint` removed in Next 16, no
eslint.config.*) — pre-existing, untouched here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-31 00:42:37 -04:00
builtbykev cdb01f97a2 The ledger id is a UUID, and only an authenticated request could have found it
`ledger_entries.id` is a uuid. The ancestry route parsed it with
Number.parseInt, so every real row would have returned 400 "invalid ledger id" —
the endpoint had never successfully served anything.

The unauthenticated probe that "proved the route was live" returned 401 from
requireAuth BEFORE the handler ran, so it could not have seen this. A
route-existence check and an acceptance test are not the same evidence, which is
exactly why the acceptance step demands a real authenticated 200 against a real
row rather than a 401.

Fixed to a UUID match, and the test asserts the real production id shape passes
while '1' and a traversal string do not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-31 00:29:22 -04:00
builtbykev e5b20b0509 Software watches coverage now, and ancestry becomes a product contract
Three closeout items, traced before building.

THE OBSERVER WAS DIAGNOSTICS, NOT MONITORING. Traced from the deployed tree:
exactly two callsites, both manual internal routes, and nothing in the scheduler
or ops path consumed it. A persistent lineage failure could have sat unnoticed
until a human asked.

The monitor now runs on the scheduler's per-minute tick, throttled to 30
minutes. It is placed there rather than after a snapshot on purpose: an
in-snapshot audit structurally cannot report that no snapshot ran, which is the
failure mode that matters most, and it would run under peak write contention
where the audit already demonstrably times out. It reads only the durable
observer and never `attachLineage`'s counters, and every failure path is
swallowed — a monitor that can take down the pipeline it watches is worse than
no monitor.

HEALTHY IS SILENCE; EVERYTHING ELSE SPEAKS. `coverageAlarm` is pure, so the
policy is testable and cannot drift into the scheduler. AUDIT_UNAVAILABLE says
"could not be measured — the audit did not run", deliberately worded so it can
never be read as "coverage is zero": those are different claims and collapsing
them is how a monitor starts lying in the reassuring direction. Alerts dedupe on
(health, cohort) so a standing fault states itself once and a NEW cohort with
the same fault speaks again.

ANCESTRY BECOMES A PRODUCT CONTRACT. It was internal-only. `GET
/api/ancestry/ledger/:id` (requireAuth, rate-limited) plus the Next proxy that
makes it browser-reachable, keyed on the LEDGER ROW ID — a stable identifier the
ledger API already returns — rather than a raw natural key exposed because it
was convenient. Its own router, so `routes/ledger.js` stays free of lineage
entirely and the grade-badge guard keeps its teeth. Every response declares
`authority: LOCAL, authority_scope: ANCESTRY_ONLY`.

THREE TEETH CAME BACK GREEN AND ALL THREE WERE MY TESTS, NOT SAFE DEFECTS.
The badge guard was CASE-SENSITIVE, so `LINEAGE_ANCESTRY` and `readAncestry`
walked straight past it. `try/finally` is valid JavaScript, so removing the
monitor's catch produced no load error and nothing asserted the containment.
And the multi-date cohort check was a grep for `.gte('game_date'` that the
head query satisfied on its own. All three replaced with behavioural tests,
including a scheduler double whose fake client HONOURS its filters — a
pass-through would have made a narrowed cohort walk look correct.

Suite 398/5,522/0 · tsc 0 · web build 0 · teeth 14/14 and 15/15.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 23:44:46 -04:00
builtbykev 1728ce4fac A cohort is a snapshot_id, not a date slice of one
Found by cross-checking the observer against an independent SQL derivation of
the same natural cohort instead of accepting its own answer.

The 2026-08-31 03:00Z run held 1,970 rows across TWO game_dates — 1,572 on
08-31 and 398 on 08-30, the late slot spanning midnight ET. The observer
selected the head row's date and filtered to it, so it measured 646 of the 801
eligible keys and reported HEALTHY. Both slices happened to be fully covered,
so the verdict was right by luck; a date slice with zero lineage would have
been invisible behind it.

The audit now covers the whole snapshot_id and reports every game_date it
spans. A test builds the real two-date shape with one slice uncovered and
requires PARTIAL_COVERAGE.

Suite 397/5,499/0 · 15/15 teeth.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 23:10:41 -04:00
builtbykev 1bdc3baec7 A valid chain can still have a hole in it
Measured against real production history while validating the ancestry
surface: `aj ewing / total_bases / over / 0.5` on 2026-08-28 carries a valid
ORIGIN from the 18:10 cohort AND a row from the 19:03 cohort that was stamped
for publication and never recorded an action.

The chain is genuinely valid and the partial row is correctly excluded from it
— that part of the contract held. But answering LINEAGE_AVAILABLE and nothing
else presents an incomplete chronology as a complete one, which is the exact
failure this surface exists to make visible, one grain finer than a whole Read.

The response now carries `chronology_complete` and `unrecorded_publications`.
The state is unchanged, because the chain that exists really is valid; what
changes is that the gap is disclosed instead of smoothed over.

Suite 397/5,498/0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 22:19:50 -04:00
builtbykev 930d526b01 Lineage productization: a live graph, an observer that doesn't trust it, and a surface that owns only what it knows
The authority review found there was nothing to switch: lineage is a
write-only graph with no living permanent writer and no product consumer.
Authority theater would have been a switch on a consumer that does not exist,
reading a store that is not being written. So: make the graph live, measure it
from outside, and expose the one category it alone owns.

WRITE-PATH FAILURE SEMANTICS, TRACED FIRST. persist(:857) -> the authoritative
cacheSet snapshot:latest(:1404) -> the lineage gate(:1426) -> ledger(:1473).
Lineage failure cannot fail base retention (earlier, separate upsert), cannot
fail publication (the product write precedes the gate), cannot fail the Ledger
(lineageIndex defaults null; the ledger has its own guard), and cannot create a
new partial row (blankLineage() first, atomicity sweep after the catch). The
isolation this tranche needed already existed; only a mode was missing.

THREE MODES, ONE EVALUATOR. `lineageWriteMode` distinguishes OFF /
CANARY_LEASED / PERSISTENT_SHADOW, and the write gate and the status surface
both read it, so they cannot disagree. The canary is CONSULTED, never
converted: its <=4h absolute expiry, its dynamic evaluation and its
fail-closed parse are untouched.

AN AMBIGUOUS CONFIGURATION FAILS CLOSED. If a sport is named by both persistent
mode and an active lease, the two instructions disagree about WHEN WRITING
STOPS — the lease says 22:45, persistent says never. The dangerous reading is
the quiet one: an operator sets a bounded lease believing writing will stop
while persistent keeps it going. We cannot know which they meant, so that sport
writes nothing until the configuration says one thing. The sport allowlist is
the canary's own, so persistent mode can never widen past it.

THE OBSERVER MAY NOT ASK THE WRITER HOW IT DID. settleLedger returned
{settled:0,pending:0} — byte-identical to a healthy "nothing to settle" — while
1,444 rows sat unprocessed, and the watchdog believed it. So `lineageCoverage`
reads durable retained state only, and THE DENOMINATOR MAY NOT CONSULT
lineage_action: eligibility is "the row was published AND a natural key is
derivable from its own identity columns", neither of which the lineage path
writes. If expectation were derived from whether lineage exists, coverage would
be 100% by construction and the metric would be decoration. Zero-expected and
zero-written are kept as different answers.

THE LEGACY BOUNDARY IS OBSERVED, NOT DECLARED. `publication_id` is stamped only
by commitPublication, so the row itself says whether lineage ran. Verified on
production: 5,353 rows carry it — 4,234 complete actions plus exactly the 1,119
historical partial rows — and zero actions exist without one. No epoch constant
is invented; a date would have been a guess about when the writer was on.
publication_id NULL -> LEGACY_UNVERIFIED. Stamped but incomplete ->
LINEAGE_UNAVAILABLE, which is the honest answer for the 1,119 and is never
quietly rewritten as legacy.

THE GRADE-SHIFT BADGE IS UNTOUCHED. revised_from_grade answers "did the letter
change"; lineage answers "which published claim superseded which". Different
questions, and a test now fails if either route learns the word lineage.

Suite 397/5,496/0 · tsc 0 · 15/15 teeth.

TWO OF MY OWN TESTS WERE VACUOUS AND A TOOTH FOUND IT. Tooth 2 came back green
because the isolation tests asserted the slate was published — true whether or
not the exception propagated — while never reaching the lineage gate at all:
the fake grader never fired `onGraded`, so the collector stayed empty and
persistedRows stayed null. Fixed by firing the hook and counting the commit.
A green teeth run means the test is missing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 22:10:14 -04:00
builtbykev 4aca33deb6 One human, one semantic identity — MLB participant convergence
The collision autopsy left two unrepaired defects, running in OPPOSITE
directions, and `outbound_collision_count` can only ever see one of them.

UNDER-COLLAPSE. Dedupe keys on `mlb:<personId>` when the participant is
proven and on the RAW PROVIDER SPELLING when it is not. Mickey Gasper
(681508) is on Boston's 40-man and not on its active roster, so an
active-only index could not identify him and every book's spelling of him
survived dedupe as its own proposition — retention was the first layer to
notice, far too late, and could only discard the loser.

SPLIT. The mirror image, and invisible to the collision metric because it
makes MORE identities, not fewer: Leo Jiménez (677870) is published as both
"Leo Jiménez" and "Leonardo Jimenez", so one human became two semantic
players in one game. Measured across the 15 MLB cohort slices since the
canonical-participant repair, this is a recurring class, not one case:
cam/cameron smith (5 slices), mitch/mitchell bratt, zac/zachary thornton,
leo/leonardo jimenez.

THE REPAIR READS MLB'S OWN RECORD. `hydrate=person` on the roster call the
pipeline already makes returns firstName / useName / useLastName, so the
legitimate name forms for a human come from the league rather than from an
alias table. An alias table is a list of the mistakes we happened to notice.
`nickName` is DELIBERATELY EXCLUDED: over 821 people it produced 14
ambiguous keys, because MLB's nickname field carries bare surnames and
shared clubhouse names — `nameKey('Smitty Smith')` is one string for both
Burch Smith and Will Smith. The four forms kept produce ZERO ambiguity.

Canonical participant reach widens to the 40-man; TEAM EVIDENCE still reads
the ACTIVE roster alone, so event admission and the impossible-binding
refusal are unchanged. Identity still fails closed: a name matching more
than one person in the event resolves to nobody.

CONTINUITY, MEASURED BEFORE WRITING ANY CODE. Over the real 19:00 cohort,
208 of 209 player_keys are unchanged and the one that moves is the defect —
`leonardo jimenez` converging onto `leo jimenez`, a key that already exists.
No new lineage family. The natural key contains game_date, so chains never
span dates and a forward change cannot fork a closed one.

DETERMINISTIC REPRESENTATIVE. Which book's payload survives was decided by
position. It is now decided by the existing MODEL_BOOKS declaration order —
reused, not authored; inventing a sportsbook ranking to settle a tiebreak
would be a market judgement smuggled in as a bug fix — with book name and a
content tiebreak. Stable under every input permutation.

TWO GUARDS, BOTH DIRECTIONS. split (one person, many identities) and merge
(one identity, many people). A merge is refused at the same single admission
seam event identity already uses; a split is counted and alerted but does not
cut the board, because it duplicates an identity rather than asserting a
falsehood.

RETENTION REMAINS AN INDEPENDENT CHECK. The old assertion grepped the source
for `player_key: nameKey(player)`. That expression stood in for a PROPERTY,
and a grep verifies a spelling. Replaced with the property itself, asserted
in both modes: when the producer emits two rows for one human, retention
still files them under one identity and still reports the collision.

Replay of the real cohort through the repair: 3,129 offerings, 100%
participants resolved, every one of 207 participants on exactly ONE semantic
key, collision 0, split 0, merge 0.

Suite 396/5,455/0 · tsc 0 · 15/15 teeth. Tooth 12 came back green first
time and that was a coverage hole, not a safe defect: nothing asserted
retention's append-only upsert. It does now.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 20:37:01 -04:00
builtbykev cc5bdf5797 A canary is temporary: bound activation with a fail-closed lease
`LINEAGE_CANARY_SPORTS=mlb` was parsed once at module load and frozen, so an
enabled canary stayed writable for the whole process lifetime. On 2026-08-29 it
was left set and the 22:00Z, 01:00Z and 03:00Z scheduled snapshots each wrote
lineage overnight with nobody watching. That history happened to be correct.
Lineage is append-only evidence, so a defective canary would have written
irreversible WRONG evidence exactly as quietly. Correct history was luck, not a
safety property.

NEW CONTRACT:  LINEAGE_CANARY_SPORTS=mlb@2026-08-30T18:00:00Z

Absolute UTC instant only -- no duration, no local timezone. A relative "4h"
would silently restart on every redeploy, which is the defect being removed.

LEGACY `mlb` NO LONGER ACTIVATES ANYTHING. It is INVALID_MISSING_EXPIRY. Leaving
it working would have left the defect in place behind a nicer-looking
alternative, so this is the load-bearing half of the repair.

EXPIRY IS EVALUATED AT THE WRITE GATE, not at startup. `isEnabled(sport, now)`
re-reads the clock on every call and `lineageCanaryEnabled` threads it through,
so a lease turns itself off with no operator, no restart, no Redis and no
network. A design where an expired canary keeps writing until someone restarts
is the same failure in a different hat.

MAX_LEASE is FOUR HOURS, and the bound is proven rather than chosen. MLB ticks
are [14,19,22,1,3] UTC; exhaustively over a minute grid across UTC date
boundaries, the shortest span enclosing THREE consecutive ticks is
22:00 -> 01:00 -> 03:00 = five hours. Four hours encloses at most two, with an
hour of margin. (An earlier note claimed six hours admitted two. It admits
three; that claim was false and the test now pins the arithmetic.) The bound is
a function of the scheduler, so a test reads the hours from sportCadence and
fails if a cadence change invalidates the proof.

Everything fails closed: absent, empty, bare sport, comma list, two leases,
duplicate @, non-leasable sport, missing timezone, numeric offset, malformed,
over-long, already-expired, expiry-equals-now, and an unreadable clock.
Configured-but-EXPIRED stays distinguishable from never-configured so automatic
containment is auditable.

No Redis, no Supabase, no counter, no network in the lease path -- an
unreachable dependency must never decide whether lineage may write.

Lineage algorithms, cache-date, participant and retention identity untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-30 14:58:11 -04:00
builtbykev 836b8c73d5 MLB odds cache is keyed on the baseball date, not the UTC date
A cache entry answers "what is the market for THIS SLATE". For MLB the slate is
a BASEBALL DATE in America/New_York -- the same date gameBinder,
retentionService, ledgerService and read_natural_key all use. The key was built
from the UTC calendar date, so between 00:00Z and Eastern midnight the key
advanced while the slate did not:

  2026-08-29T01:03Z = 2026-08-28 21:03 ET
    slate date 2026-08-28,  key looked up  odds:mlb:2026-08-29

A value written earlier that evening under odds:mlb:2026-08-28 was then
unreachable -- not expired, ADDRESSED WRONG.

CAUSAL HONESTY: this is NOT retroactively the cause of the failed 01:03Z canary.
That run also used an 84-minute-old observation whose cache had passed its 1h
TTL. Two independent reasons; the date defect is real but not proven
counterfactual.

MLB ONLY. Every other sport keeps the UTC basis -- their date semantics are
unproven here and rekeying a cache they already write and read consistently
would invalidate live entries for no demonstrated defect.

Symmetry is structural, not conventional: the three readers that built the key
inline now ask `oddsService.getCacheKey(sport)`, so writer and readers cannot
diverge. The ET date comes from `scheduleService.gameDateET` -- the repository's
own Intl/America\/New_York primitive, now exported -- so DST is the zone
database's business and never offset arithmetic. An unresolvable clock REFUSES
rather than falling back to the other basis.

TTL truth is kept separate: a correctly addressed but expired entry still
misses, and CACHE_TTL is unchanged at 3600.

Provider priority, quota policy, retries, EARLY_RETURN_ODDS_ERROR semantics,
lineage lookup, game-date repair, canonical participant and intraday belief
integrity are all untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-29 14:42:10 -04:00
builtbykev 7efb04e280 Lineage family lookup: bound it to a slate, and say what an action is
TWO DEFECTS, one lookup.

SCALE. The family lookup sent 100 natural keys as a PostgREST IN-list.
`read_natural_key` has NO pg_stats row at all -- the table's last autoanalyze
(2026-08-26) predates the column ever being populated -- so the planner used a
default per-value selectivity, estimated 172,409 rows and chose a sequential
scan of 344,818: 8.5s, then 57014. At 50 keys the same shape returned in
~357ms. The cliff is a statistics artifact, not a volume one, which is why the
repair does not depend on the estimate improving and is not CH=50.

`readNaturalKey` builds `sport|game_date|player_key|stat|side|line[|#event]`,
so SPORT AND GAME_DATE ARE COMPONENTS OF THE KEY. Two rows sharing a key
necessarily share both, and scoping the lookup to the (sport, game_date) pairs
present in the requested keys is LOSSLESS BY CONSTRUCTION. One index-backed
range per date, walked with safePaginate; cost is bounded by ONE SLATE however
long the chronology gets. Measured: 5,000 keys -> 1 scope, and the plan is
`Index Scan using model_snapshots_lineage_family_idx, cost 0.28..1.92`.

VALIDITY. A row carrying `read_natural_key` is not history: the key is stamped
on every candidate BEFORE the lookup, so a failure leaves it on a row that
never became an action. Proven this was not cosmetic -- fed the raw rows the
old lookup returned, the resolver produced a REVISION with a NULL read_id (an
orphaned chain node) and labelled a brand-new Read LEGACY_UNVERIFIED.
`isValidLineageAction` states what a completed action IS: all nine fields, in
the query and again in code.

ATOMICITY. A failed attempt now leaves NO lineage-specific state.
`publication_id`/`published_at` are untouched -- the slate really was
published, and erasing a true fact to tidy a false one is the wrong repair.

Replayed the exact failed 19:00Z cohort through the real resolver, side-effect
free: 119 NEW / 379 CHANGED / 621 UNCHANGED -> ORIGIN 119 / REVISION 379 /
RECAPTURE 621, 0 wrong parent, 0 wrong ordinal, 0 null read_id, 0 forks --
byte-identical with all 1,119 failed partial rows present. Clean-head parity
1,024/1,024.

Migration 050 is CONCURRENTLY + IF NOT EXISTS, drops nothing, rewrites nothing.
Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 19:56:56 -04:00
builtbykev 889a96584a Intraday is a market overlay: it may not restate belief
`grade` is a band label over `p_win` and nothing else. `servedGrade.gradeFor`
reads exactly { p_win, refused, refusal_reason, insufficient_data,
factor_adjustment } -- no line, no odds, no edge, no side -- and
`gradeFreeze.TOLERANCES.MAX_LETTER_MISMATCH` is 0, so the served letter MUST
equal the band of the served probability.

The intraday refresh re-graded an adversely-moved prop AT THE CURRENT LINE and
then copied only `res.grade` onto the row for the LOCKED line, discarding the
p_win that produced it. Two defects in one write: the letter stopped matching
the probability beside it, and the letter answered a different proposition than
the claim it was stamped on.

MEASURED on ledger_entries in the current model era
(engine1@2026-08-07-fullwindow): 30,047 unrevised rows carry ZERO letter/p_win
mismatches; all 9 intraday-revised rows are mismatched. The revision was the
sole producer of incoherent authoritative rows. The pre-cutover era is
uninterpretable here (its `grade` was the engine index, not a p_win band, and
mismatches 49.4% of the time WITHOUT any revision) -- which is exactly why the
control was run before quoting a rate.

Simulated over the 57 historical revisions: in the current era, keeping the
published grade restores coherence on 9 of 9.

So the adverse branch now records the move and stops there. No grader is
invoked, no ledger revision is applied, and the grade-rank comparator is
deleted rather than parked -- a dead comparator beside the code is how the
mutation gets re-wired. A production coherence gate compares every output row
to its input on every BELIEF field and, on any divergence, publishes the
untouched slate instead.

History is untouched: the 57 existing `revised_from_grade` rows keep their
values, and the UI that renders them is unchanged. Forward-only.

Lineage is not touched and stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 18:48:07 -04:00
builtbykev f7cc19772b Canonical MLB participant: prove the human, then dedupe
VYNDR had three definitions of "same player": raw p.player (dedupe),
utils/normalize.normalizeName (features, NO nicknames), utils/playerName.nameKey
(retention, WITH nicknames). Retention was the first layer to notice they
disagreed, and all it could do was discard the loser — 28 collisions.

The fix is not a better spelling. The event-scoped roster ALREADY carries the
StatsAPI personId; buildPlayerTeamIndex was fetching and discarding it.

SCOPE IS LOAD-BEARING, and this is the finding that shaped the design. Measured
on the real 2026-08-28 league rosters (1,583 rows, 30 teams): a bare name key is
NOT globally unique — `max muncy` (ATH 691777 / LAD 571970), `jose fermin`
(665877/820862) and `luis garcia` (472610/671277) each resolve to TWO different
humans. Team-scoped: 0 ambiguous. Event-scoped across 33 real events including
the verified doubleheader date: 0 ambiguous. So resolution is scoped to the two
teams actually playing, and FAILS CLOSED — no event, no game, no candidate, or
more than one candidate yields null and the prop keeps prior behaviour.

FEATURE IDENTITY PARITY, measured on both real cohorts before writing the patch
(504 participants):
  RAW_SUCCESS/CANONICAL_SUCCESS/SAME id      365
  RAW_SUCCESS/CANONICAL_SUCCESS/DIFFERENT id   0   <- required 0
  RAW_SUCCESS/CANONICAL_FAIL                   0   <- required 0
  CANONICAL_AMBIGUOUS                          0   <- required 0
  RAW_FAIL/CANONICAL_SUCCESS                   7   <- repair, not regression
The seven improvements are exactly the alias class: Michael->Mickey Gasper,
AJ->A.J. Ewing, JT->J.T. Realmuto, Mike->Michael Busch, Richard->Richie
Palacios. This is why raw spelling could not be trusted: player_id_map holds
`mickey gasper` and NOT `michael gasper`, so under the raw name the lookup
outcome depended on which alias happened to survive dedupe. Provider arrival
order was deciding model input availability.

ALL SEVEN COLLISION GROUPS PROVEN SAME_PLAYER_ALIAS against the authoritative
roster for the exact event date — one personId each (681715, 699625, 676356,
695491, 673357, 681508, 671739), zero false normalizations, zero unresolved —
and all seven collapse to one participant under the new key, so the 14 colliding
propositions merge upstream instead of being discarded downstream.

FALSE-MERGE SIMULATION on both cohorts: cross-event 0, different-stat 0,
different-line 0, false merges 0. Note honestly: the retained rows cannot
exhibit the merges themselves, because retention already discarded the losers —
so the merge half is proven directly on the alias groups, the safety half on the
cohorts.

RETENTION IS UNTOUCHED — schema, player_key, conflict identity and the collision
counter are all unchanged. collision_count must reach 0 by upstream repair, never
by making the counter lenient. A teeth proof injects raw spelling into the
retention identity and the suite rejects it.

Source provenance preserved: p.player is never overwritten; the participant
rides beside it as mlb_person_id + canonical_player_name.

Twelve teeth against a green baseline of 117 — dedupe back to raw name (3),
feature lookup back to the alias (2), resolver stops blocking ambiguity (1),
lookup degraded to raw (1), event scope dropped (1), event/line/stat dropped
from the key (2/2/2), provenance destroyed (1), raw spelling in retention
identity (2), collision_count accepted nonzero (2), and a guard proving the
frozen game-date repair cannot be reverted (10). Restored byte-identically.
Teeth #4 (canonical resolving to a DIFFERENT feature id) is covered by the
real-data parity measurement rather than an injected branch, because it is a
data comparison, not a code path — stated plainly rather than claimed as a test.

No hardcoded player names in production code — a test greps for all seven.

4 files, 86 insertions. retentionService, gameBinder, oddsService,
acquisitionTrace, probabilityEstimator, bookRoles, playerName, normalize and
snapshotScheduler: UNCHANGED. Admission rules, MODEL_BOOK filters and model
formulas: 0 changed lines. The game-date repair is intact.

391 suites / 5,339 tests pass. web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 12:54:35 -04:00
builtbykev 7d579fd8c1 Repair the MLB timing contract at its producer seam
A prop carrying a usable game_time skipped bindGame — and bindGame is where
game_date was assigned. So the fast path left props TIMED and DATELESS. Once
PropLine began supplying game_time (historically it did not, which is why the
binder exists), every prop took that branch.

PROVEN ON PRODUCTION, attempt acq_972544a1-563e-47ba-bde4-2f73ecf536bd:
12,015 props acquired live, alreadyHad 12,015, dates_requested 0, schedule games
0, unresolved 12,015 (no_schedule), admitted 0, graded 0, no product writes.

THE DERIVATION IS NOT INVENTED HERE. bindGame already defines exactly this for
the already-timed case — `game_date: etDate(prop.game_time)` — and
ledgerService.gameDateFor already PREFERS the derived ET date over a
provider-supplied one (`dateET(prop.game_time) || prop.game_date`). etDate and
dateET are the same America/New_York en-CA formatter. This applies the contract
the fast path skipped; it does not add a new one.

Repaired at the PRODUCER, not by teaching consumers to compensate. The prop
already has an exact game_time, so the date is derived deterministically — the
props are NOT sent back through event rebinding to obtain it. Canonical event
resolution remains solely responsible for proving WHICH game, which is what
matters for doubleheaders, provider nesting defects and team contradictions.

SCOPE, locked by tests:
  valid game_time + missing game_date   -> filled
  valid game_time + matching game_date  -> preserved
  valid game_time + DIFFERING game_date -> preserved, NOT silently rewritten
                                           (future-hardening territory)
  invalid/absent game_time              -> no date invented; existing binder
                                           fallback semantics unchanged
  game_time itself                      -> never modified

UTC DATE IS NOT THE BASEBALL DATE, and the tests prove it: 01:45Z -> 2026-08-27,
03:10Z -> 2026-08-27, plus a DST-boundary pair (2026-11-01 01:30Z -> 10-31 EDT,
06:30Z -> 11-01 EST). No hardcoded offset; the repository helper does the work.

DOUBLEHEADER: both halves derive the SAME game_date — expected, and what makes
the schedule fetch possible — while the UNMODIFIED resolver still separates them
by clock into gamePk 824514 / 824478.

Ten teeth, injections verified present, against a green baseline of 129:
assignment removed (9) · UTC date instead of ET (5) · already-timed props forced
through rebinding (4) · conflicting date overwritten (2) · invalid time given a
fabricated date (1) · doubleheader guard weakened (1) · unresolved admitted (6) ·
game_time modified on the bound path (2). Restored byte-identically.

TWO TEETH LANDED AND PASSED FIRST TIME — coverage holes, not safe defects. The
doubleheader fixture matched both games exactly, so deleting the >1h guard
changed nothing; and no test reached the bound path at all. Added a stray-time
refusal case and a bound-path case, then re-ran both failing.

Admission still fails closed: UNRESOLVED / AMBIGUOUS / CONTRADICTED, and RESOLVED
without a canonical id, are all still rejected. The gate protected production
during the outage and is untouched.

ONE FILE, SIX BEHAVIOURAL LINES. eventIdentity, gradeSlateService,
snapshotService, retentionService, oddsService, bookRoles, analyzeViaEngine1,
probabilityEstimator, acquisitionTrace and snapshotScheduler all UNCHANGED.

390 suites / 5,310 tests pass. web tsc exit 0. Lineage stays OFF.

The historical 03:00 cause remains UNPROVEN — that run was served from cache and
its game_date state was never recorded. Consistency is not proof.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 03:15:58 -04:00
builtbykev f54b0627e1 Truthful provenance for an operator-invoked snapshot
The internal one-shot route already called the SAME production runSnapshot with
the SAME default dependencies — but it passed no trigger, and runSnapshot
defaults an absent trigger to SCHEDULED. So every operator-invoked run was
recorded as though the cron had fired it. That was a lie about provenance,
present by omission, and it would have contaminated the trace of any forced
diagnostic run.

CONTROLLED_FORCED is now its own trigger. Both internal routes (`/snapshot/:sport`
and `/snapshot/all`) stamp it, along with the process generation. The scheduler
still stamps SCHEDULED, and a test asserts neither internal route can label
itself scheduled.

Trace retention moves from "scheduled only" to a named allow-list of SCHEDULED +
CONTROLLED_FORCED. INTRADAY is still refused — it runs every ~20 minutes and
would displace scheduled evidence, which is the failure the store exists to
prevent. MANUAL_API stays refused too.

THE PIPELINE IS UNTOUCHED. snapshotService, gradeSlateService, retentionService,
snapshotScheduler, oddsService and eventIdentity are all UNCHANGED. A test
asserts the route injects no dependency override — no getOdds, gradeAndCacheSlate,
retention, ledger, cacheSet/cacheGet, gameBinder, eventIdentity, mlbAdapter or
notify — so the only difference from a scheduled invocation is the label and the
absence of a scheduled hour, which a forced run genuinely does not have.

The ?limit bisect-hook invariant is preserved and tightened: the opts passed
carry exactly {trigger, processStartedAt} and never a stray limit.

Teeth, injections verified present, against a green baseline:
  :sport route mislabelled SCHEDULED -> 3 fail
  /all route mislabelled SCHEDULED   -> 2 fail
  intraday admitted to the store     -> 5 fail
  trigger filter removed             -> 3 fail

THE FIRST TEETH RUN WAS INVALID AND IS DISCARDED: both routes live in one file,
so a single-occurrence replace hit `/snapshot/all` and left `/snapshot/:sport`
correct — the injection landed on the wrong target and the suite passed. Coverage
for `/all` was added, plus a test that the file contains exactly two
CONTROLLED_FORCED stamps and zero SCHEDULED ones, then both were re-run failing
independently.

Four stale assertions updated with the reason recorded: three pinned the
`not_scheduled` refusal string (now trigger-agnostic) and one pinned an empty
opts object on the route.

389 suites / 5,280 tests pass. web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 02:56:22 -04:00
builtbykev c1d9ec5bbb Path coverage: every branch from acquisition to the first grade callback
Audited the corridor rather than trusting the recorder. Nine upstream operations
sit between acquisition success and gradeAndCacheSlate — game binding, per-date
schedule fetch, roster index, attachEventIdentity, book-price capture, team-stat
refresh, hits-factor context, matchup keys — each with its own catch. NONE of
them early-returns, so the corridor always reaches the grader; but seven of them
were SILENT, and two could be misattributed.

COVERAGE WAS INCOMPLETE. Closed:
  * BINDING — bound / unresolved / already-had, and its throw
  * SCHEDULE — dates requested, game count, and its throw. A schedule outage
    previously surfaced as an EVENT IDENTITY error because it lands in that
    catch; it now records SCHEDULE_STAGE_ERROR and rethrows unchanged.
  * ROSTER — indexed players/teams/failed and evidence-date-validity, and its
    throw. Its catch only console.warn'd, so this stage was entirely invisible.
  * DEDUPE per-reason accounting off the filter's OWN branches: invalid fields /
    non-model book / duplicate identity / capped / not examined, plus
    model-book eligibility. Counts reconcile to the input exactly.
  * GRADE BOUNDARY — `reached` is derived from a candidate count and proves
    nothing. GRADE_LOOP_ENTERED, FIRST_GRADEBESTSIDE_STARTED and
    FIRST_ONGRADED_OBSERVED are now separate control-flow facts. No model
    output is recorded; a test greps for p_win/grade/confidence/edge/side.

Terminal states now name the stage: SCHEDULE_STAGE_ERROR, ROSTER_STAGE_ERROR,
BINDING_STAGE_ERROR, EVENT_IDENTITY_STAGE_ERROR, ADMISSION_STAGE_ERROR,
DEDUPE_STAGE_ERROR, IDENTITY_ALL_UNRESOLVED, ALL_REJECTED,
DEDUPE_ALL_NON_MODEL_BOOK, DEDUPE_ALL_INVALID_FIELDS, DEDUPE_EMPTY_OTHER,
READY_FOR_GRADING, GRADE_LOOP_STARTED, FIRST_GRADE_CALLBACK_OBSERVED.

TRACE COMPLETENESS INVARIANT. `reconcilePregrade` — acquisition NONZERO +
CONTINUED with no correlated downstream state is an OBSERVABILITY_GAP, never a
pipeline verdict. This programme has twice read an absence as a conclusion
("MLB exited at acquisition", "all props rejected at admission"); both were
wrong. Now it is a typed state with tests.

dedupeProps takes an OPTIONAL stats object and increments on the branches it
already takes, in the same order — reused, never reimplemented. Without the
object it is byte-identical; a test asserts that.

A PRODUCTION-BREAKING BUG CAUGHT BY THE FULL SUITE: the frozen no-op recorder
did not implement gradeStarted/firstOnGraded, so any caller without a recorder
threw inside the grade loop — and gradeAndCacheSlate's catch turned that into
{written:false,count:0}. Every slate would have graded NOTHING, silently. Fixed,
NO_PREGRADE now covers the full recorder surface, and a test asserts it does.

Exception semantics unchanged throughout: every added catch records and RETHROWS
the identical error. Admission rules, dedupe predicates, MODEL_BOOKS, event
identity, gameBinder, gradeBestSide and its arguments: 0 changed lines.
eventIdentity, oddsService, retentionService, gameBinder, bookRoles,
analyzeViaEngine1, probabilityEstimator, snapshotScheduler and ledgerService:
UNCHANGED. Zero new external calls — the only diff hit is the existing
getScheduleWithPitchers line re-indented into its own try.

Twelve teeth, injections verified present, against a green baseline of 89:
upstream catch silent (1) · missing trace as failure (1) · admission exception
as ALL_REJECTED (2) · non-model-book as duplicate (3) · dedupe-empty as
rejection (1) · falsely says grading started (6) · onGraded unrecorded (1) ·
sport overwrite (2) · intraday overwrite (1) · different attempt id (3) ·
observer adds an external call (1) · store failure changes outcome (2).
Restored byte-identically; teeth 3/4/6 re-run after the NO_PREGRADE fix.

The first teeth pass ran against a baseline the finer states had invalidated;
six superseded assertions were updated first and the run repeated.

388 suites / 5,269 tests pass. web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 01:21:09 -04:00
builtbykev 8d741e3932 Pre-grading stage trace: record which branch loses the cohort
The acquisition recorder proved the previous diagnosis wrong: the 03:00 MLB
attempt acquired 6,365 props from a propline cache hit and CONTINUED. Zero
reached the first grading callback, and nothing durable says why.

WHY REPLAY WAS IMPOSSIBLE. The 03:00 input was the cache object written
02:20:24.889Z. ODDS_CACHE_TTL defaults to 3600s, so it expired ~03:20; it is
now gone. No historical copy exists: closing_captures holds no MLB rows past
01:40:52, model_snapshots none, and `bookprices:mlb` carries no timestamp the
book-comparison route exposes. Calling /api/odds/mlb was REFUSED — a cold cache
would fetch and fire recordDownstream -> gradeAndCacheSlate, writing product
state and manufacturing a false recovery. So the input is NOT RECOVERABLE and
no replay was attempted.

A SECOND CLAIM IS WITHDRAWN. The previous tranche concluded "all 6,365 props
were rejected at event admission", reasoning that dedupe cannot empty a
non-empty list. That is FALSE: `dedupeProps` also FILTERS — it drops props with
no player/stat_type/line and any prop whose book is not a MODEL book. A fully
admitted cohort can still dedupe to zero. Demonstrated in test. So the failing
branch was never established, only assumed — which is exactly what the order
forbade, and why DEDUPE_EMPTY is a first-class outcome here.

THE REGION IS SMALL AND EVERY BRANCH LOOKS IDENTICAL FROM OUTSIDE:
  identity annotation (runSnapshot, mlb only, own catch)
  -> admitForGrading (pure, can throw)
  -> dedupeProps     (pure, can throw, ALSO filters)
  -> mapLimit(gradeBestSide)   <- first onGraded-capable call
gradeAndCacheSlate swallows every throw in that region and returns the same
{written:false,count:0} it returns for an honest zero.

The recorder distinguishes them: READY_FOR_GRADING · ALL_REJECTED ·
IDENTITY_STAGE_ERROR · ADMISSION_STAGE_ERROR · DEDUPE_STAGE_ERROR ·
DEDUPE_EMPTY · NO_INPUT_PROPS · OTHER_PREGRADING_ERROR. An exception is never
folded into ALL_REJECTED, and with no admission evidence the classifier refuses
to classify at all — a test pins that.

EXCEPTION SEMANTICS UNCHANGED. Each stage is wrapped to record and then RETHROW
the identical error, so the enclosing best-effort catch still handles it exactly
as before: nothing caught that was not caught, nothing swallowed that was not
swallowed. Admission and dedupe rules, the impossible-binding invariant, event
identity, model books and the first-row-wins cap are untouched — the only
behavioural lines in the diff are `const gate/unique` becoming `let`.

Correlated to the SAME snapshot_attempt_id the acquisition recorder minted — not
a new run id — and stored under its own key `ops:pregrade:{sport}` so it can
never displace the acquisition record. Same atomic LPUSH/LTRIM pattern, bounded
per sport, scheduled-only, best-effort at the call site, auto-disabled under
test.

CAUGHT DURING BUILD: `savePgTrace` was defined and NEVER CALLED — the recorder
would have persisted nothing, the same "built, correct, never invoked" failure
this programme has hit before. A test now drives the real runSnapshot and
asserts a trace is persisted carrying the acquisition's attempt id.

Ten teeth, injections verified present, against a green baseline:
empty collector read as ALL_REJECTED (2) · admission throw reported as
admitted=0 (1) · dedupe throw reported as output=0 (1) · dedupe-to-zero
mislabeled as rejection (1) · identity exception hidden (1) · observer reorders
the candidate array (1) · trace failure changes product outcome (1) · intraday
overwrites scheduled (1) · different attempt id (2) · secret leak (1).
Restored byte-identically.

THREE OF THOSE LANDED AND PASSED FIRST TIME — coverage holes, not safe defects:
the dedupe-throw and identity-throw tests only exercised the recorder directly,
never the real path, and nothing asserted the CANDIDATE array is not reordered.
All three closed with real-path tests, then re-run failing.

Two stale source assertions updated with the reason recorded: both pinned the
exact `const gate = …` / opts-key order that the recording wrapper changed.

387 suites / 5,237 tests pass. web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 00:41:34 -04:00
builtbykev 238f0f67cf Scheduled acquisition trace: record the fork instead of inferring it
MLB stopped producing anything at the 22:00 and 01:00 slots on 2026-08-27 while
NBA/WNBA ran normally. The differential narrowed it to exactly two runSnapshot
exits — getOdds THREW, or getOdds RETURNED ZERO PROPS — and production retained
nothing able to tell them apart. A failed acquisition has no snapshot_id, writes
no ledger row and updates no slate, so it left no durable evidence at all. The
alert channel could not fill the gap either: quota/test-alert reports sent:true
while the ntfy topic replays 0 messages, so absence of alerts is not evidence.

This records the decisions the existing control flow already makes. It changes
no acquisition behaviour: provider order, the single existing retry, the quota
threshold, the fallback and cache policy are all untouched. The only behavioural
line in the diff is getOdds gaining an optional recorder argument, defaulted to
a frozen no-op so every existing caller is byte-identical.

IT MAKES NO PROVIDER CALL. Measured: zero added fetchAllOdds/getProps/gateway
calls, and the trace module contains no HTTP of any kind. Tests assert both.

WHY REDIS, NOT MEMORY. server.js arms the scheduler in EVERY process and there
is no lock or leader election, and a rolling deploy demonstrably serves two
containers at once — so process-local evidence could be written by a container
nobody later probes. Storage is LPUSH + LTRIM, which is atomic: two schedulers
racing the same slot both survive instead of one silently overwriting the other,
and each attempt carries a process_generation so they stay distinguishable.

WHY A BOUNDED HISTORY, NOT "LAST ACQUISITION". The scheduled MLB attempt failed
at the hour and an intraday attempt SUCCEEDED ~20 minutes later. A single
last-value would have erased the failure with the success — precisely the
evidence needed. Only SCHEDULED attempts are retained (persist refuses any other
trigger), each sport keeps its own key, and intraday structurally cannot write
one because intradayRefreshService never calls runSnapshot.

WHAT IS CAPTURED, per attempt: cache decision; PropLine outcome as
NONZERO/ZERO/ERROR/NOT_ATTEMPTED with count; the odds-api fallback with
allowed_at_invocation, blocked_reason and the quota AS OBSERVED AT THAT
INVOCATION — reading provider quota hours later and calling it historical
evidence is the exact mistake this exists to stop; then the final result and the
runSnapshot terminal outcome. The EXISTING retry appears as a second attempt; no
retry was added.

Sanitized: keys, tokens, URLs and long opaque strings are redacted, and no prop
payload is retained — counts only. Tests assert a dirty provider error and a
real prop array both come out clean.

BEST-EFFORT AT THE CALL SITE, not just in the default dep — a teeth proof showed
an injected store could still throw into a healthy snapshot. Now any
implementation is safe. persist also auto-disables under NODE_ENV=test unless a
client is injected (the opsNotify precedent); without that the default path
opened a real ioredis connection inside every suite driving runSnapshot.

Read-only GET /api/internal/acquisition/:sport behind the existing internal
auth. It runs no pipeline and makes no fetch.

Nine teeth against a green baseline, injections verified present:
intraday overwrites scheduled (1) · one global slot (2) · thrown getOdds with no
terminal trace (2) · zero mislabeled as error (1) · quota not captured at
invocation (1) · observer makes a provider call (1) · credential leak (2) ·
scheduled/intraday share an identity (1) · telemetry failure breaks the snapshot
(1). Restored byte-identically.

TWO OF THOSE LANDED AND PASSED FIRST TIME — coverage holes, not safe defects:
the provider-call scan did not forbid getOdds, and nothing exercised a
non-scheduled trigger through runSnapshot. Both closed, then re-run failing.

Model, retention, event identity, admission, dedupe, ledger, lineage, cadence,
quota tracker and the PropLine adapter are all UNCHANGED. Lineage stays OFF.

386 suites / 5,210 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 22:29:42 -04:00
builtbykev 048e4eaa3f Event-aware retention identity: two games, two receipts
One player prop in Game 1 and the same-looking prop in Game 2 are two different
historical claims. The retention conflict identity did not know that.

SEMANTIC IDENTITY FIRST. Two outbound rows are the same retention proposition
within one cycle when they share the cycle, the EVENT, the participant, the
stat, the line and the side. Book is deliberately absent — collapsing books is
dedupeProps's actual job and the price anchor is chosen later. The database
index is enforcement of that answer, never the definition of it.

THE EVENT COMPONENT NEVER FABRICATES. canonical_event_id where a sport has a
resolver — MLB's admission gate rejects unresolved/ambiguous/contradicted props
BEFORE grading, so every row that can reach retention has one — and game_id
otherwise, which is NOT NULL in the schema and is the only event label sports
without a resolver possess. Both are in the identity, so the weaker label still
discriminates where the stronger is absent.

NULLS NOT DISTINCT IS LOAD-BEARING, NOT STYLISTIC. canonical_event_id is NULL
for every non-MLB row. Measured on a disposable PG17: under PostgreSQL's default
semantics the same NBA proposition inserted twice produced TWO rows — every
retry duplicating for ever. With NULLS NOT DISTINCT the same test yields one.
That measurement is what rejected the plain composite option.

MIXED-FLEET BRIDGE. A rollout serves both builds at once (measured 11/12 new,
1 old). Old and new writers need different indexes and NO schema state satisfies
both: with the legacy index present a new writer fails 23505 on a doubleheader;
with it gone an old writer fails 42P10. A bare ON CONFLICT DO NOTHING would have
bridged this, and PostgREST does not emit one — `ignoreDuplicates` WITHOUT
`onConflict` was measured raising a real duplicate-key error, so that bridge does
not exist through this client.

So the writer bridges it. It targets the event-aware identity and, on exactly
the two errors meaning "the schema is not in the state I expect" (42P10, or
23505 NAMING the legacy index), retries the SAME chunk on the legacy target. A
failed chunk rolls back atomically — measured 0 rows — so the retry cannot
double-write. Correct in every schema state: legacy-only and both-present
degrade to legacy semantics with no outage; new-only keeps both games.

The bridge is deliberately narrow. A supersedes conflict is ALSO a 23505, and
swallowing it would destroy the forked-history guard, so the legacy index must
be named. All three model_snapshots writers (persist, commitPublication,
recoverFromFork) go through it; no hardcoded legacy target survives.

MEASURED, through the real supabase-js -> PostgREST -> Postgres path on
production-shaped PG17:
  * 1,000 REAL propositions from the verified 2026-08-17 STL@CIN doubleheader
    (1,738 retained rows under ONE game_id), replayed across both real gamePks:
    OLD index materialized 1,000 of 2,000 — 1,000 LOST. NEW index materialized
    2,000 of 2,000 — 0 lost.
  * retry idempotency, over/under, line, stat, player, non-MLB same-game and
    non-MLB different-game all behave correctly under the new index.
  * ORDINARY-SLATE PARITY over ALL 434 real cohorts / 328,262 retained rows:
    old identities 328,262, new identities 328,262, delta 0, cohorts changed 0.
    The index is therefore guaranteed creatable and nothing historical splits.

CONFLICT_IDENTITY is now DERIVED from RETENTION_CONFLICT rather than restated —
a test caught them silently disagreeing, which is exactly how the materialization
check could have expected an identity the database no longer enforced.

EXPAND/CONTRACT are separate files on purpose. 048 is additive and retires
nothing; 049 drops the legacy index and must not be applied until fleet
convergence is proven by sampling, never assumed from a fast rollout.

NO BACKFILL. Legacy rows keep NULL canonical_event_id and remain LEGACY
EVENT-AGNOSTIC RETENTION, which is what that NULL truthfully says.

The materialization defence is untouched and now reports the bridge honestly:
while the legacy index still collapses a doubleheader, expected 4 vs actual 2
yields MATERIALIZATION_MISSING and the cohort is refused.

Nine teeth, injections verified present, against a green baseline of 97:
1 event distinction removed (10) · 2 phases collapsed (2) · 3 bridge swallows
everything (6) · 4 NULLS NOT DISTINCT removed (1) · 5 old-container error as
success (3) · 6 semantic/DB identity disagree (8) · 7 collision detector removed
(2) · 8 partial transport usable (3) · 9 collision unannounced (1).
Restored byte-identically.

Model and product untouched: gradeSlateService (event-aware dedupe), event
identity, ledger, calibration, chain, lineage config and the status route all
UNCHANGED. Zero cacheSet changes, zero web paths, schema contract unchanged (no
new columns). Lineage stays OFF.

385 suites / 5,178 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 20:30:45 -04:00
builtbykev 11277a1b99 Materialization truth: a completed write is not a complete cohort
Transport truth says every intended write succeeded. Materialization truth says
every identity that should exist actually exists. The retention writer could
only report the first, and the gap is not theoretical.

THE CONFLICT IDENTITY, traced to the real index:
  model_snapshots_cycle_prop_uniq UNIQUE (snapshot_id, player_key, stat, line, side)
written by `upsert(..., { onConflict: same five columns, ignoreDuplicates: true })`.
Measured in production: no key column is ever NULL (0 of 328,262 rows), so
NULLS DISTINCT never applies and the identity is plain column equality. `line`
is an unconstrained numeric, so identity normalises it — 0.5 in and "0.50" out
must not read as two identities for one stored row.

PRIOR-CYCLE COLLISION IS IMPOSSIBLE. `snapshot_id` is in the identity and is a
fresh UUID per cycle, so no row can be suppressed by an earlier cycle.
Append-only chronology across cycles is safe, and `captured_at` is not in the
identity, so a cohort cannot be split by timing.

INTRA-CYCLE COLLISION IS REAL, AND WE CAUSED IT. `canonical_event_id` is NOT in
the identity. A doubleheader — same hitter, same stat, same line, two genuinely
different games — is ONE identity. Demonstrated through the real collector: 4
outbound rows, 2 distinct identities, 2 rows discarded by ignoreDuplicates with
no error, `written` counting all 4 and the terminal status reading COMPLETE.
Before event-aware dedupe the second game was dropped before grading, so the
collision could not arise; that fix moved the loss downstream into retention.

The conflict identity is NOT changed here — that is a separate decision with its
own before/after. This makes the loss visible instead of silent.

EXPECTED vs ACTUAL. `expectedMaterialization(rows)` derives the identity set
from the FINAL outbound payload using the exact database identity — never from
`attempted`, which counts rows sent, not identities that can exist.
`reconcileMaterialization` compares SETS, not counts: two sets of equal size can
still differ, and a cohort that swapped one identity for another passes every
count test ever written. A collision passes set equality by construction (the
discarded row was never in the expected set) while real rows were lost, so
collision_count > 0 fails the cohort on its own.

A cohort is evidence-complete only when transport is COMPLETE, missing = 0,
extra = 0, and collisions = 0.

OBSERVABILITY stayed minimal. `last_retention` was already PER SPORT (a Map
keyed by sport), so no fix was needed there and the route is UNCHANGED — the new
fields ride the existing entry: outbound_rows, expected_materialized_count,
outbound_collision_count, expected_identity_digest. Counts and a digest only,
never the identities, which carry player names. The expected set is the one
materialization fact unrecoverable from the database afterwards, which is why it
is the only thing recorded at runtime.

A collision leaves transport COMPLETE, so the existing failure alert could never
see it. It now has its own high-severity alert naming the counts, the cycle and
the build, and says the cohort is not evidence-complete.

Seven teeth, each injection verified present, against a GREEN baseline of 63:
  1 attempted===written as evidence completeness   -> 1 fail
  2 COUNT(*) equality instead of set equality       -> 1 fail
  3 snapshot_id dropped from expected identity      -> 4 fail
  4 unexpected collision allowed to qualify         -> 1 fail
  5 single global last_retention slot               -> 2 fail
  6 partial chunk failure treated as usable         -> 3 fail
  7 collision loses its announcement                -> 1 fail
Restored byte-identically (retention 742f116473d97f49, snapshot 81129facbabeb280).

Three brittle assertions repaired, with the reason recorded: two windowed on a
byte count that a neighbouring block outgrew — a test failing because of its
neighbour, not its subject — now windowed to syntactic landmarks; and one
counted TERMINAL.COMPLETE occurrences, which a legitimate comparison
incremented. It now asserts one DECISION and one READ.

persist() and createCollector are BYTE-IDENTICAL. onConflict and
ignoreDuplicates appear in the diff only as prose. Model, event, ledger,
calibration, chain, lineage config, and the status route: UNCHANGED. Zero
lineage/publication files, zero cacheSet changes, zero web paths. Lineage OFF.

Schema contract unchanged: release 64, prod 67, prod-only 3 (debt, not
authorized), missing in prod 0.

384 suites / 5,144 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 19:44:56 -04:00
builtbykev 9809626c99 Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.

So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.

TERMINAL RETENTION STATES (retentionService.classifyPersist):
  NOTHING_TO_PERSIST       attempted 0 — a refusal-only slate is still a cycle
  SKIPPED_NO_DATABASE      no database configured; not a failure
  COMPLETE                 attempted > 0, written === attempted, no error
  FAILED_ZERO_WRITE        written === 0 — first chunk failed
  FAILED_PARTIAL           0 < written < attempted — a later chunk failed
  FAILED_UNRESOLVED_ERROR  counts look complete but an error is unresolved;
                           unreachable through today's loop, and kept because
                           the alternative is reporting COMPLETE holding an error

The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.

classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to 35da190.

`written` counts rows in COMMITTED CHUNKS, not database inserts: the upsert uses
ignoreDuplicates, so a re-run legitimately inserts far fewer rows than it writes.
Comparing written to count(*) will disagree by design. Documented, because that
mismatch is exactly what would be misread as a partial write.

VISIBILITY. The 35da190 alert condition was
`r.error || (!r.skipped && r.attempted > 0 && r.written === 0)` — it could not
see a partial cohort as a distinct state. It is now driven by terminal status,
so FAILED_PARTIAL alerts as loudly as a total failure and is labelled INCOMPLETE
and unusable as evidence. Best-effort is unchanged: the product continues and
the alert says so.

OBSERVABILITY. A successful cycle previously left only a console.log with no
snapshot_id, no code_sha and no terminal status, so completion could not be
established after the fact. `GET /api/internal/snapshot/status` now returns
`last_retention` per sport — sport, snapshot_id, attempted, written, status,
completed_at, code_sha, error_summary — taken verbatim from the persistence
result. Existing internal auth, read-only, counts and status only, no payloads.
No new table, no new route.

RELEASE-AUTHORIZED INSERT CONTRACT. The migration-derived contract is the
release authority; production is not. A prod-only column is DRIFT / RECORDED
DEBT and never becomes permission by existing. Verifier classifies: release
column missing in prod -> HARD FAILURE; prod-only -> drift warning; outbound key
outside the contract -> contract failure (enforced against the real upsert
payload). It is read-only and never rewrites the contract from live schema.
Live: release 64, prod 67, prod-only 3, missing in prod 0.

Six teeth, each with the injection verified present, against a green baseline:
  1 written>0 as generic success        -> 6 fail
  2 later-chunk failure reports COMPLETE -> 5 fail
  3 FAILED_PARTIAL does not alert        -> 3 fail
  4 status reports a recalculated count  -> 1 fail
  5 row presence treated as completion   -> 1 fail
  6 invalid outbound column reintroduced -> 4 fail
Restored byte-identically (retention b341cf16c1baa992, snapshot 81ab1bd7730dee89).

Two stale assertions updated rather than deleted, with the mechanism change
recorded: the alert-shape tests described the superseded written===0 condition,
and the runtime probe test pinned an exact import list.

Model and product preserved: analyzeViaEngine1, probabilityEstimator,
gradeSlateService, lineageCanaryConfig, eventIdentity, ledgerService,
calibration and chain all UNCHANGED; zero lineage/publication files touched;
zero cacheSet changes; zero web paths. Lineage stays OFF.

383 suites / 5,118 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 19:12:59 -04:00
builtbykev 35da190f2c Retention hotfix: drop published_side, derive the schema contract, break the silence
`createCollector.onPublished` set `published_side` beside `published`.
`published_side` is not a model_snapshots column. supabase-js declares the
UNION of row keys in the `columns=` parameter, so one invalid key made
PostgREST reject the ENTIRE batch with a 400 — every sport, every cycle.
Retention is best-effort, so nothing surfaced. Confirmed in edge logs.

The field was redundant as well as invalid: `side` is already on the row.
Deleted rather than added to the schema — a column would preserve an
accidental artifact.

Three things missed it, and each is now closed:

1. WRONG SHAPE INSPECTED. The manual check sampled the collector after
   onGraded only and never called onPublished, so the offending key was
   not yet on the row. It read a pre-publication shape and reported the
   final outbound shape as clean. The new test captures the array actually
   handed to .upsert(), after the full production call order.

2. NO CONTRACT. Every retention test injects a permissive fake client that
   accepts any column set, so 381 suites proved the logic and never once
   compared a row against the database. The contract is now DERIVED — the
   migration chain applied to a disposable postgres, read out of
   information_schema (scripts/generate-schema-contract.js). A
   hand-maintained list would be a second opinion about the schema, and a
   second opinion is what let this through. scripts/verify-schema-contract.js
   checks the contract still describes a live database.

3. SILENT FAILURE. A failed batch reached one console.log. It now emits a
   high-severity structured event carrying sport, snapshot id, stage,
   error, code_sha and timestamp. Best-effort semantics are unchanged —
   the product continues and says so — but the failure is observable.
   `skipped` (no database configured) is not a failure and does not alert.

Teeth, each with the injection verified present before the run:
  - published_side back into the final payload -> 4 tests fail; restored
    byte-identically (sha 6a0ced7c52134135 both sides)
  - settled_at (a REAL contract column) -> accepted, so the guard
    discriminates by contract membership, not by novelty
  - alert block deleted -> 3 tests fail; restored byte-identically

Model and product behaviour untouched: analyzeViaEngine1,
probabilityEstimator, gradeSlateService, lineageCanaryConfig all unchanged.
Lineage stays OFF. Net source change is one behavioural line plus the alert.

382 suites / 5,094 tests pass. web tsc exit 0 (zero web paths touched).

Measurement blackout recorded, NOT backfilled: last good retention write
2026-08-27T19:08:32Z; ceaa896 started 21:16:41Z; the 22:00 UTC cycle ran
(ledger wrote 22:05:02) and persisted zero snapshot rows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 18:53:37 -04:00
builtbykev ceaa896f77 Runtime observability: report the build and canary state the system acts on
The rollout stalled at RUNTIME_UNVERIFIED because two facts were answerable only
as a side effect of a scheduled snapshot writing a row: which build is running,
and whether MLB lineage is effectively enabled. Every state transition therefore
waited on cron rather than on asking the service.

- src/services/lineageCanaryConfig.js — THE canary resolver. Parsed once at
  module load (unchanged semantics), normalised sorted/deduped/trimmed, frozen.
  snapshotService's write gate now delegates to it, and the status probe reads
  the SAME state. A route that parsed the environment itself would be a second
  version of the truth, free to drift from the gate it claims to report.
- GET /api/internal/snapshot/status gains runtime.code_sha (the production
  codeSha resolver — never git, never gitea/main; null when unavailable),
  runtime.started_at (computed ONCE at module load, so it marks a boundary
  rather than reading as now; deliberately not called deployed_at), and
  lineage_canary {enabled, sports, configuration_source}.
- No raw environment value is returned; sports is the normalised set and
  configuration_source says only ENVIRONMENT vs DEFAULT. Router-wide
  requireInternalAuth is unchanged: 200 with key, 401 without.
- Effective lineage config is fixed for the process lifetime, so
  runtime.started_at is a defensible lower bound for how long that state held.

Strictly observational — the handler still only reads Redis.

Model and decision code byte-identical to 8c6aef1: analyzeViaEngine1,
probabilityEstimator, gradeRanking, eventIdentity, gradeSlateService,
retentionService, ledgerService, mlbStatsAdapter.

Suite 381/5,081/0 from the release worktree; web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 17:16:05 -04:00
builtbykev 8c6aef1e12 Event admission gate: an unresolved or contradicted game does not earn a Read
A fractured identity keeps two bad records from merging. It does not make an
unknown game true. Until now a prop with an unresolved or verified-impossible
event still continued into grading under that synthetic key; it no longer does.

- event_binding_status contract: RESOLVED / UNRESOLVED / AMBIGUOUS /
  CONTRADICTED / UNSUPPORTED. CONTRADICTED and UNRESOLVED stay distinct — one
  means we know the association is wrong, the other that we do not know.
- ONE admission gate (gradeSlateService.admitForGrading), before dedupe and
  before grading. Admission requires RESOLVED *and* a canonical_event_id; a
  legacy derived game_id can never satisfy it. Rejected props are returned, not
  discarded, so retention keeps them as evidence.
- Roster evidence is now DATE-SCOPED. statsapi honours ?date= and it changes the
  answer (Joe Mack is on the 2026-08-26 Marlins roster, absent on 2026-04-15).
  Evidence that does not describe the slate's date can only yield UNRESOLVED,
  never CONTRADICTED — uncertainty must not become an accusation.
- A mis-nested market is REFUSED, never re-bound. Knowing Joe Mack is a Marlin
  does not license moving a provider record into the Marlins game; that would
  invent provenance.

Measured on the real 2026-08-26 slate: 2,878 props -> 2,874 admitted, 4 rejected
(EVENT_PLAYER_TEAM_CONTRADICTION), each player's correct game still resolving.
Doubleheader 824514/824478 both remain independently RESOLVED and admitted.

No change to probability, projection, side, grade, confidence, ranking,
normalization or calibration. Lineage remains disabled.

Suite 380/5,061/0 from the release worktree; web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 16:50:50 -04:00
builtbykev 352016790a MLB canonical event identity, impossible-binding refusal, event-aware dedupe, publication commit
Release-isolated slice built from 41ba38e. Ships ONLY the event-integrity +
publication + lineage-canary closure; the 90-path development tree stays
undeployed.

- canonical MLB event identity from statsapi gamePk (mlb:gamepk:<pk>), with
  event_identity_source/method/version recorded. The id is canonical; the
  binding is derived and says so.
- IMPOSSIBLE-BINDING REFUSAL. Verified in prod 2026-08-26: Joe Mack (Marlins)
  was bound to Dodgers@Braves and Yandy Diaz (Rays) to Rangers@WhiteSox, both
  from one book in the 01:00/03:01 UTC cycles after their own games began. Root
  cause is source market data, not the binder. A prop whose player's team is not
  an event participant now refuses; unknown team preserves uncertainty.
- event-aware dedupe: books still collapse, events no longer do. An unresolved
  MLB event fractures rather than falling back to the collision-prone
  date+teams key.
- publication commit moved AFTER the authoritative Redis slate write, with
  exact parity-gap identity when the product publishes and the record does not.
- lineage dual-write behind LINEAGE_CANARY_SPORTS, DISABLED for this deploy.

Excluded deliberately: WNBA feed/chain, market ontology, PerformanceDistribution,
calibration certification, truth diagnostics, applyRevision Phase-1, and the
analyzeViaEngine1 confidence-rounding change (a served field).

Suite 380/5,040/0 from this worktree; web tsc exit 0; champion output identical
to production.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 16:29:38 -04:00
builtbykev 6c34af3414 checkpoint: chain shadow, WNBA possession feed, baseball chain
Backup commit of uncommitted working-tree state found during Legion
recon (Tony resurrection, STEP 0). This work existed only on the
laptop disk.

- chain shadow accrual + probe script (038_chain_shadow.sql)
- WNBA possession feed: ESPN adapter, usage service, verify script
  (039_wnba_player_game.sql)
- baseball chain
- retention/snapshot service updates, tableKeys, matchupKeys
- specs: chain-v1, wnba-possession-feed, wnba-source-survey
- unit tests for the above

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QnvJAkC3h5QGmb6dipoiWn
2026-08-14 16:53:37 -04:00
builtbykev 0657b71d18 Drop lock_lines and retire its writer (dead table, dead writer)
lock_lines was built in Session 64 for a staleness audit - join lock-time
per-book lines to closing_captures and ask whether our locked line was
stale-high vs consensus. That audit was never written. ruler-comparison.sql
records why: only 43 settled rows ever joined it with >=2 two-sided books.

Measured before removal: 367,595 rows, 104 MB, zero readers in src/, scripts/
or web/src/ - the only from('lock_lines') was an upsert, every other mention a
comment. Zero dependents: no FK, no view, no trigger. The newest pg_dump held
all 367,595 rows, pg_restore-verified before the drop.

The write is off too, because a dead table that keeps refilling is only half
solved: it was accruing 36,440 rows/day, 10.3 MB/day, 23% of all database
growth, for a question nobody was asking. buildLockRows is kept and still
tested - the logic was never what was wrong, and re-arming is one flag plus
re-creating the table.

DB 510 MB -> 406 MB: 81% of the 500 MB cap, +94 MB headroom, under it for the
first time in months.

Moat and grade untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 01:18:10 -04:00
builtbykev fb0010222d Stop persisting missed_window refusals (halt the bleed at source)
missed_window means 'this capture pass ran after first pitch'. It is a fact
about our cron cadence, not the market: once a game starts the same prop emits
a fresh refusal every ~20 minutes for the rest of the night, per book, per
side. Measured over 7 days of production that is 332,608 rows/day - 83.1% of
all closing_captures writes, ~66 MB/day - and since B1 filtered both readers,
nothing reads them.

The filter lives in persist(), not buildCaptureRows(), and that is the whole
trick: the caller computes the capture-rate alarm from the full in-memory
array, so filtering at build time would have blinded the ops alarm to the exact
condition it exists to catch. captureRateAlarm is pure; a test asserts the
caller still passes the full array, and that a 10-priced/90-late pass still
fires at 0.10.

Narrow by design: one_sided_price still persists (liquidity signal), priced
captures unchanged, and the rare fault refusals still persist because each
names a pipeline fault worth seeing. A missed_window row carrying a price is
kept.

Growth drops 400,469 -> 67,862 rows/day. The 4.2M historical rows are now
static, so the cleanup is a calm decision rather than a race.

Nothing deleted. Grade untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 00:34:04 -04:00
builtbykev 2271f46ab2 dCLV close leg: read priced captures only (the unfixed twin)
computeDirectionalForRow selected the latest closing_captures row with no
missed_reason filter, while its sibling attachClosingProb has had one since the
CLV instrument repair. closing_captures records a refusal for every prop x book
x side on every cycle after first pitch, so a refusal always carries a later
captured_at than the last real price - latest-first returned a refusal on
26,448 of 26,448 identity groups, and computeDirectionalClv refuses on a
missedReason. That is why dclv_state has been 'unknown' on 100% of rows since
Session 64.

Measured on 600 real settled rows: 100% unknown becomes flat 45.8%, negative
23.0%, positive 22.8%, unknown 8.3%. ClvBadge will render MOVED TOWARD US 114
and MOVED AWAY 117 per 600 - near-symmetric, which is the honest shape.

Existing rows do not recompute (first-computation-wins). The re-stamp is
described in BUILD-STATE, not run.

The grade is untouched. No rows deleted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-12 00:14:12 -04:00
builtbykev f61ec6b391 Read integrity, as-of context, and the shadow matchup resolve (A1-A7)
Seven orders of measurement-first repair. The served grade does not move.

A0/A1 — the unordered page walk returned the right COUNT and the wrong ROWS:
410-617 of 2,490 duplicated with an equal number never returned, while
rows.length matched the server exactly. safePaginate orders on a real unique
key, verifies the tuple at runtime, and THROWS on a query error instead of
treating it as end-of-data. Both hits PROVES are withdrawn: they were drawn
through that reader, and defense_by_direction's distinct-n was likely below
the gate floor all along.

A2/A2b — rolled across every reader: 11 FAIL -> 0. Composite keys pulled from
pg_index (the context tables are dated-composite and had no single unique
column). The unordered helper is deleted, not parked.

A3 — ledgerService and retentionService defaulted the SAME env var to
DIFFERENT versions, so no ledger row ever carried the marker eligibility
requires. One source now. model_snapshots settlement moved onto the cron:
15,484 -> 28,894 settled, repaired-champion 0 -> 7,556.

A4 — hitsFactorContext takes an as-of cutoff. Refusal over reconstruction: no
row at-or-before the date means the factor does not apply, never the nearest
row. Live path unchanged, proven 400/400 on real rows.

A5 — factor_inputs freezes what the factor READ, never the multiplier, so an
audit can recompute and check. It also recorded the finding: the three hits
factors have NEVER fired. prop.opponent and prop.opposing_pitcher are read by
the resolver and written by nothing.

A6/A7 — matchupKeys resolves those keys from the posted lineup plus the
schedule's probable pitchers, and fires the factors into a SHADOW freeze:
248 fires on 308 props, 245 of which would move the grade. The served
forecast is untouched. specs/a8-shadow-factor-gate.md pre-registers the test
that decides whether they ever go live.

Nothing is turned on. CALIBRATION_DEPLOYED stays []. Both verdicts stay
withdrawn. 4,772 tests / 371 suites green, web build exit 0, read-integrity
harness 34/34.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-11 22:49:56 -04:00
builtbykev 71d3b7b786 E10 Report issue template + E12 /report archive, to spec
PHASE 0 — the spec, read not recalled. E10: "Hybrid: dark billboard header
that survives every client, light paper body Gmail can't wreck. 600px,
stacked, no webfont dependence." Content law: "One email per slate day.
Top read, what changed, the record. Nothing else." E12: "EVERY ISSUE SHOWS
ITS OWN DAY RECORD -- THE ARCHIVE IS A LEDGER TOO."

COMPOSED, NOT FORKED. The audit had E10 as PARTIAL, not absent:
newsletterService already builds the daily report's CONTENT and lints its
voice. What was missing is the designed hybrid SHELL, so reportTemplate.js
is a template over that builder rather than a second report -- the same
call made for the movement strip, and for the same reason.

PHASE 1 — the hybrid shell is an ENGINEERING constraint, not a look, and
the tests say so: Gmail strips style blocks, Outlook ignores flexbox, and
a dark body renders as a black rectangle in several clients. Hence tables,
inline styles, 600px fixed, system fonts, no image required to read, and
the green SHIFTS from #00D4A0 to #00A57D on paper because the dark-mode
green is unreadable there.

FACT-CONTRACTED: a section whose data is absent is OMITTED and NAMED in
`omitted`, never filled. There is no code path producing a placeholder
figure. The honesty block carries the real numbers -- graded count,
cleared-ceiling count, the realized rate against baseline, and that we do
not issue A grades.

E1'S LAW TRAVELS EVEN THOUGH ITS RENDERING CANNOT. An SVG strip is not
reliable in email, so movementText carries the RULE: green only when the
move favours the read, and a flat market says FLAT · [N]D rather than
showing nothing.

NO DESIGNER SAMPLE DATA. Nabers 1,120.5, No 128, DAY RECORD 9-4 are a spec
for what a live issue renders; pasting them in would be fabrication
carrying a designer's authority and would look entirely correct. Tested.

PHASE 2 — /report is now the real archive, REPLACING the S41 redirect to
/blog. That redirect existed because the surface did not; E12 built it, so
the placeholder is correctly gone and the S41 test is updated rather than
worked around. Every row carries its own day record, and an unknown record
says UNSETTLED -- never a dash that reads as zero. Empty archive is an
honest state.

Backend: public read-only /api/report over Redis issues, plus the Next
proxy. Both surfaces registered under the reachability guard.

A test bug I made twice now: my check for forbidden sample values matched
the template's own doc block, which NAMES those values as things never to
paste. Documentation worth keeping, so both suites strip comments before
matching -- a guard that reads its own warning is not reading the code.

WAVE-2 STATUS: E1, F9-F11, E10, E12 done. Still gated -- F5 article media
and E16/F8 on the card-system reconciliation; the in-season hub IA on the
social chat's formula; E9/E15 on model; E2/E6 on licensing.

Read-only throughout; serving fingerprint unchanged including
newsletterService; accrual clock unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 22:17:59 -04:00
builtbykev 49e76068da Doctrine + E1 movement strip + F9-F11 offseason hub shell, to spec
PHASE 0 — specs/ARCHETYPE-TAXONOMY-DOCTRINE.md records the ruling as
shared law: 83 designed glyphs are the full four-sport taxonomy; a glyph
renders ONLY where its archetype is modeled and proven. 39 of 83 map to a
real archetype and are wired; the 44 unmapped are DORMANT SLOTS for
WNBA/NBA/Soccer, not a wiring gap. Wiring them would mean inventing 44
archetypes to consume artwork -- decoration presented as classification,
which is forbidden. DUAL THREAT and PAINT BOSS are modeled archetypes with
no mark: the mirror gap, flagged to the design side. When a sport's
archetypes ship, activation is a MANIFEST lookup, not new art.

PHASE 1 — E1 movement strip. The spec's own line is "the movement strip is
defined once here and reused everywhere a line has a past", so it is a
primitive, not a fourth chart.

RECONCILED RATHER THAN FORKED: lib/gradeShift.js ALREADY implements E1's
colour law -- toward/against/flat, including the direction flip that makes
an UNDER's favourable move the opposite sign of an OVER's. MovementStrip
CONSUMES buildGradeTimeline instead of reimplementing it, and a test
asserts it never redefines isUnder. GradeShift stays the grade-history
view; this is the reusable strip. That is the card-fork lesson applied
before it could happen again.

Spec laws honoured: STEPS NOT CURVES (H then V, no smoothing -- a curve
invents prices that never traded, and a test rejects any C/S/Q/T command);
green only when the move FAVOURS the read; FLAT renders as a hairline plus
FLAT · [N]D because a flat market is a finding; and too little history
says NO MOVEMENT HISTORY rather than rendering blank.

PHASE 2 — F9-F11 offseason hub shell, built from Vyndr Offseason.dc.html.
The spec's load-bearing words are used verbatim: "OUTLOOKS REPRICE ON NEWS
· NOT GAME ODDS" (an offseason number is not a game line), the QUIET WIRE
empty state ("No outlook-moving news since X. We don't manufacture
movement."), WHAT CHANGED TODAY as the hero with the countdown ambient and
top-right, the tag-colour-is-meaning row anatomy, the open -> NOW -> FAIR
triplet with the movement strip embedded, and the OUTLOOK ONLY block where
every row carries NOT GRADED.

THE DESIGN FILE'S SAMPLE DATA IS NOT IN THE COMPONENT. Wembanyama +420 ->
+330, Nabers cleared 11:42 AM, the Summer League names -- all of it is a
SPEC for what a live feed renders, and copying it in would be fabrication
carrying a designer's authority. A test asserts none of those strings
appear.

The IN-SEASON information architecture is NOT invented here. The spec
covers an offseason hub; nothing specifies how content, articles, wire and
the live slate share year-round navigation. That remains the open design
gap, and the route notes it.

Two test bugs caught and fixed: my first assertions matched my own doc
comments -- the ordering check found "WHAT CHANGED TODAY" in the header
block and the no-curves check caught the word "curve" in the sentence
explaining why curves are wrong. A guard that reads its own explanation is
not reading the render; both now strip comments first.

PHASE 3 — both surfaces registered under the reachability guard. Read-only
throughout, serving fingerprint unchanged, accrual clock unchanged at 0
eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 21:24:36 -04:00
builtbykev 60469422af Wave D1 primitives — and the audit says E1/E9 were not Wave 1
PHASE 0 — the order proposed E1 + E9 as Wave-1 and told me to follow the
audit if it disagreed. It disagrees: E9 calibration curve is WAVE D3,
gated on MODEL work ("resolve n>=20 vs N30, accrue buckets"), and E1
movement strip is WAVE D6, a large surface build. E9's gate is live right
now -- calibration is WITHDRAWN at 0 eligible dates, so the curve could
only render its empty state today. Building it would ship a component
whose entire purpose is unavailable.

PHASE 1 — three of the five real Wave D1 items were ALREADY DONE, and the
2026-07-31 audit has aged:

  D1 glyph library  audit: 38/83 wired (46%)
                    now:   COMPLETE for everything wireable -- 39 of 83
                           designed glyphs map to a real archetype, all 39
                           are wired, colours match the registry exactly
                           (0 disagreements).
  A1 card token     audit: BUILT-BUT-DRIFTED, "in only 1 file"
                    now:   BUILT-TO-SPEC -- it IS the --bg-1 token,
                           consumed by 32 files. The audit counted literal
                           hex, which is what a correctly tokenised value
                           looks like.
  B1 boundary blue  audit: PARTIAL, hex in 2 files
                    now:   BUILT -- --priced-out/#8fb2de is a token with a
                           documented colour law, 4 consumers.

The 44 unwired glyphs are NOT a wiring gap: they have no backend
archetype, so wiring them means inventing 44 archetypes to consume
artwork -- the fabrication this programme refuses. That is the 41-vs-74
scope question and it is Kev's call. Separately, 2 registry archetypes
have NO designed glyph (DUAL THREAT, PAINT BOSS) -- a design gap.

PHASE 2 — what was genuinely absent is now built. web/src/lib/motion.js:
nudge() capped at 180ms so it reads as acknowledgement rather than
latency; bootStagger capped at 240ms because uncapped, row 40 waits 1.1s
and the stagger BECOMES the latency it exists to disguise; rowHover
returns handlers not CSS so touch cannot stick a hover state; and
revealOnIntersect returns an unobserve in every path and reveals
IMMEDIATELY when there is no IntersectionObserver or motion is reduced --
content is never hidden behind a capability check.

Reduced motion is honoured, not softened. The sharpest of the 10 tests:
bootStagger under reduced motion returns opacity 1, not merely delay 0 --
if the CSS animation supplies the opacity, skipping it leaves the row
invisible forever.

PHASE 3 — the reachability guard gains a PRIMITIVES section: a module
built to be embedded must declare its exports AND name its intended
consumers, because a primitive imported by nothing is the same
built-but-unread class as an unmounted component.

WAVE-2 UNGATED: F9-F11 offseason hub, F5 article media, E10/E12 Report,
E1 movement strip. GATED: E9 + E15 on model, E16/F8 on the resolution tail
and the card-system reconciliation, E2/E6 on licensing, E13 on another
order.

Read-only throughout; serving fingerprint unchanged; accrual clock
unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 20:31:55 -04:00
builtbykev c575a708c7 Content studio API + preview page; widen the reachability guard; correct
two inventory errors

INVENTORY CORRECTION, and it was mine. Phase 2's two "orphans" are NOT
orphans -- my board grepped only web/src/app and missed component-level
mounting. The transitive check says both are already mounted:

  BookComparisonPanel -> GradeResultCard -> app/scan/page.tsx
  NewsWire            -> ExploreHub      -> app/explore/page.tsx

So book comparison is DONE (wired to /api/books, rendering on the grade
card) and THE WIRE is DONE-BY-DESIGN, mounted in ExploreHub. Its header
names an "Offseason Hub" as its home, and that hub genuinely does not
exist -- but that is board item #8, not a mounting bug, and inventing a
surface to satisfy a comment would be the wrong fix.

The lesson is the same one this session keeps teaching: I checked one
directory and reported a conclusion the check could not support.

ALSO CAUGHT: I overwrote src/routes/content.js, which was the Session-29
content-templates route, by picking a filename without looking. Restored
from git with no work lost; the new surface lives at
/api/content-studio and both now coexist.

PHASE 0/1 — /api/content-studio serves finished posts (copy, branded card,
card_svg, the fact_contract each was REQUIRED to have, and the facts that
actually backed it) plus a POST for editorial status in Redis. Private via
internal key; the Next proxy holds the key server-side so the browser
never does. /studio renders it as a thin client -- copy and card side by
side with the fact contract visible, because reviewing copy by reading it
is exactly how a wrong number ships. Never-blank: a night with nothing
generated says so.

API-FIRST is the point: the endpoint an autonomous poster will call is the
one the page already renders, so the agent handoff is a pointer change,
not a rebuild. Contract documented at docs/CONTENT-STUDIO-API.md.

EXPRESS 5 BROKE 23 SUITES at first: `router.get('/:date?')` throws at
mount time in Express 5, taking down everything that imports app.js. Two
explicit routes instead.

PHASE 3 — the reachability guard is widened from grade-fields-only to a
general built-but-unread check. Book comparison, THE WIRE and the content
studio are now registered surfaces; a page counts as its own entry point
(Next mounts it by convention) while everything else must trace to one.
22 checks green; a registered-but-unimported surface still goes red.

FULLY ISOLATED: read-only on model/slate/ledger, serving fingerprint
verified unchanged, accrual clock unchanged at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 18:51:52 -04:00
builtbykev 74aa75945e Content engine: posts that structurally cannot lie
PHASE 0 — contentEngine makes Truth Law structural, not careful. Copy is
token-substituted and an unbacked {token} REFUSES to render -- there is no
code path that produces a plausible default. The fact contract is asserted
before any string is built. Card and copy render from ONE fact object, so
a caption and a card cannot disagree. No live model writes factual claims:
the voice is in the template, the facts are pulled, and the voice-polish
port is deliberately unwired, because an LLM that can rewrite a sentence
can rewrite a number.

18 tests carry the proof. The one that matters most: ZERO IS PRESENT.
"0 cleared B+" is our most honest possible post, and treating 0 as missing
would be the Number(null)===0 breach wearing its opposite coat -- it would
silently delete exactly the post the brand is built on.

PHASE 1 — three templates, generating real posts from tonight's data:
hot hitters off the repaired full-season log, the honesty flex off the
real servedGrade distribution (2,140 graded / 70 cleared B+ / 42% not
separable / A unissuable), and streaks verified from settled outcomes only.

THE ENGINE CAUGHT A BUG IN ITSELF, and it is the sharpest lesson here. The
first run published "No hitter is meaningfully hot tonight -- we could
dress up a middling week as a streak. We don't." That was FALSE: the
box-score cache spans only the settled window, every player had under 20
games, and the pool was empty. A broken pull was publishing as considered
editorial judgement -- the fourth appearance of this class tonight and the
first where our OWN HONESTY COPY was the disguise.

Fixed structurally rather than by patching the number: an absent() variant
may now DECLINE to speak, and the template separates "no candidates at
all" (SKIP with a reason) from "candidates judged, none hot" (honest
absence). Both locked by test. Source corrected to mlbStatsAdapter.fullLog,
the same log the repaired champion reads.

PHASE 2 — cardRenderer emits SVG rather than canvas: it is text, so it
diffs in review and its numbers are greppable, which matters when the
whole claim is that the numbers are real. VYND white + R green, slashed-Y,
scanlines, mono. The card never formats its own facts -- every string
arrives pre-rendered and gate-checked.

PHASE 3 — scripts/generate-content.js writes copy + card per template to
.content-out/<date>/. Template N+1 is a registry entry: requires, pull,
copy, card, absent. Queued as stubs, not built: hot takes, daily reads,
"grades we DIDN'T give", cross-sport streak variants (the streak template
is already sport-agnostic -- settled outcomes and a noun).

FULLY ISOLATED: read-only on every source, zero writes to serving, model
or ledger tables. Serving fingerprint verified unchanged. The accrual clock
is untouched at 0 eligible dates.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 16:34:57 -04:00
builtbykev 55b210cb95 Fix the dormant basketball window-bug before it ships; guard the class
PHASE 0 — audit. espnStatsAdapter's slice(0,20) was already fixed at
494c83c, so the ESPN branch of getStatRows inherits a full log and is
CORRECT. MLB's l20_avg is a real seasonTotal/games aggregate, verified,
also CORRECT. TWO basketball defects remained:

  DEFECTIVE  nbaGameLogFeatures   m20 = avg(vals.slice(0, 20))
             -- l20_avg is the season reference projectionFor reads, and
             slicing to 20 made it a twenty-game average wearing a season
             label. The MLB l20 bug, unfixed for basketball.
  DEFECTIVE  getStatRows python   getGameLogs(playerName, sp, 20)

PHASE 1 — both fixed via a named SEASON_LOG_DEPTH = 100, past an 82-game
season so a request can never truncate one. API COST: ZERO. The count is a
request parameter, so asking for a season is the same single call. No
extra request, no extra quota.

MEASUREMENT DEFERRED, EXPLICITLY: basketball is offline, there are no
settled basketball rows, and resolution before/after CANNOT be measured
now. This is a code fix, not a measured improvement -- exactly like the
deferred MLB sibling paths.

PHASE 2 — the window guard asserts the property on source: no fixed N may
stand in for a season. It immediately caught TWO MORE instances I had
missed in Phase 1 -- a hardcoded 20 at featureCache:377 and
gameLogService's own `count = 20` DEFAULT, which would have handed a
twenty-game window to any caller that omitted the argument. That is a
seventh path, found by the guard rather than by me.

Retro-confirmed: run unchanged against 981a05c it goes 3 failed / 6
passed, flagging the basketball slice, the game-log request and the
season-reference check.

The class in full, now six paths across two sports, every one of which
looked like ordinary code. `slice(0, 20)` is unremarkable; what made it a
defect was the QUESTION it answered -- "what is this player's season
rate?" -- and no test could see that mismatch because the value produced
was always a plausible number.

PHASE 3 — THIS CLOSES THE NON-ACCRUAL ARC. Everything buildable without
settled rows is built: push unblocked (it was the wrong remote, not the
firewall), grade surface honest and rendering, reachability guarded,
champion repaired across every path including dormant basketball, and the
window class guarded so it cannot return.

The program is now correctly IDLE on modeling. Today's count: 0 eligible
dates, all four re-audit items WAITING. FIRST TRIGGER: 10 eligible
calibration dates on the repaired champion, at which point the resumption
order is calibration re-fit, hits factor lift, prior verdict re-audit,
rbi lineup-slot gate.

No basketball measurement claimed. No NBA chain/archetype build. MLB
serving verified byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 15:00:24 -04:00
builtbykev 981a05cbd6 Render-reachability guard: make built-but-unread a CI failure
Three consecutive orders shipped a backend-correct field that never
reached a screen, all on a green suite: gradeBands (required by no
serving code), served_grade (dropped at the adapter boundary),
GradeScaleLegend (imported by nothing). Each was caught by luck on a later
re-check, and in two of the three I had already reported the wiring done.

WHY GREEN TESTS COULD NOT SEE IT: backend tests stop at the API payload.
They prove a field is PRODUCED and say nothing about whether it is
CONSUMED. Invisible by construction, not an oversight in any one test.

THE TRAP, NAMED: the difficulty pools in the backend, so by the time a
field exists on the payload it feels finished. What remains is a
three-line adapter change nobody considers worth verifying, so it gets
claimed rather than traced. The last inch is the one with no friction,
which is exactly why it gets skipped. "I added the field" and "a user can
see it" are different claims and only the first is fun.

THE GUARD traces each promised field the whole way: payload -> adapter
consumes -> component renders -> component is MOUNTED. Mounted is
transitive to a Next entry point (page/layout/template), the only thing
that puts a pixel on screen, depth-limited so an import cycle cannot hang
the suite.

Container rows are exempted EXPLICITLY, not silently: served_grade carries
container:true plus a rendersVia list, and a separate assertion checks
every named part actually renders. The exemption is auditable and cannot
hide an unrendered field.

The guard tests itself -- an orphan component must report unmounted, and
the contract must be non-empty, since an empty contract passing vacuously
is how this would most plausibly rot.

RETRO-PROOF: run unchanged against 3591c76, before the wiring, it goes
11 failed / 8 passed and names the exact bugs -- "the ceiling stance /
grade scale legend - its component is MOUNTED, not merely written", "the
served grade object - the adapter consumes it", "whether the band
separates from the baseline - a component actually renders it". Green on
the current tree.

HONEST SCOPE LIMIT: gradeBands is NOT in the contract and would not be
caught. It is a backend module, not a promised user-facing field, and it
is correctly unwired -- every band collapses to base-rate at current
resolution. Out of scope by design, not oversight.

Now in the standing suite, so the three-gate floor is tests green
(including reachability) + build exit 0 + fingerprint. No serving or model
change. p_win never mutated. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:44:17 -04:00
builtbykev 3591c7626e Total grade cutover + the ceiling stated as a position
PHASE 0 caught my own repeat of the failure I diagnosed one order ago.
91927a4 attached `served_grade` BESIDE the old letter and left `grade`
alone -- so the honest grade reached nobody, exactly as gradeBands had
been built-correct-and-unread. grep showed served_grade appearing in one
file (where I set it) and all 14+ consumers -- scan route, dashboard,
parlay, newsletter, desk, content templates, retention -- still reading
`.grade`, i.e. still the dishonest letter.

CUTOVER IS NOW TOTAL: legacy.grade IS the honest letter. Overwriting the
one field every consumer already reads cuts every surface over at once
instead of editing fourteen call sites and missing one. engine1's index is
preserved as `engine_grade` and verified read by ZERO serving code.

Confidence follows the letter: it came from a grade-band midpoint of the
OLD letter, so leaving it would have paired a served B+ with a C's
confidence. Both now derive from p_win, kept on the existing 0-100 scale.

MEASURED BLAST RADIUS before shipping: 303 of 47,991 non-refused
snapshots (0.6%) have a grade but no p_win, and now render NO READ instead
of a letter. That is correct -- their old letter came from the retired
index carrying 0.48% resolution, i.e. noise -- and NO READ is a rendered
state with a reason, so never-blank holds.

PHASE 1 — the ceiling is now a STATED POSITION, not a confusing absence.
servedGrade.SCALE_LEGEND plus web GradeScaleLegend.tsx say it plainly: we
do not issue A grades, no band has hit at a rate that would justify one,
our honest ceiling is a strong B+ (~66% realized vs ~60% baseline), and if
the model earns an A the legend changes and we say why. The
separates_from_base_rate flag renders per band -- C+/C/C- are labelled
"we cannot separate this from the baseline", which is most of any slate.

PHASE 3 hand-verified across every state: B+ with 3 factors (basis
forecast_plus_matchup_factors), B+ with none (forecast_only), C flagged
not-separable, F, and three refusal states rendering NO READ with reasons.
never-blank PASS, no-manufactured-A PASS.

Test fallout was real and is documented rather than papered over: engine
BEHAVIOUR assertions moved to engine_grade, suppression assertions stayed
on grade (a suppressed prop has no letter either way), and the confidence
78 -> 95 change is the grade-band midpoint being replaced by p_win.

No A-threshold loosening. No calibrated number leaks (deployed set empty).
p_win never mutated. Ten frozen modules verified unchanged including
engine1 and probabilityEstimator. No Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 14:07:02 -04:00
builtbykev 91927a4a8a Serve an honest grade: the letter was carrying 1/6 the information of the
number beside it

PHASE 0 corrects the order's premise. A grade letter has been served all
along -- engine1.gradeProp builds it from an additive factor index,
computed INDEPENDENTLY of p_win. gradeBands is orphaned for a different
reason than assumed: it defines what a letter MEANS from realized
outcomes, and every band collapses to base-rate at current resolution.

The measurement that changed this order, on 3,417 settled props:

  grade  n      realized  mean p_win
  A         8    0.500      0.647     <- the TOP grade did WORST
  B       985    0.640      0.700
  C     1,695    0.602      0.676
  D       303    0.558      0.604
  F       426    0.535      0.588

  letter resolution 0.00116 (0.48% of variance)
  p_win  resolution 0.00715 (2.98%)
  -> the letter carried 0.16x the information of the number beside it

Concretely, from the hand-verify: Christian Encarnacion's 0.95 over
graded C and his 0.05 under ALSO graded C -- same hitter, opposite
forecasts, same letter. The gap was never that grades don't ship; it is
that the weaker of two available signals shipped as the headline.

PHASE 1 — model/servedGrade.js derives the letter from p_win with bands
anchored on MEASURED realized rates (B+ 0.663 / B 0.646 / C+ 0.615 /
C 0.589 / C- 0.548 / D 0.512 / F 0.447, base 0.6005).

NO MANUFACTURED A, structurally: A+/A/A- are UNISSUABLE, not rare. The
realized rate plateaus at 0.65-0.68 above p_win 0.70, so no band has
earned a top letter; a test sweeps every p_win 0..1 and asserts none
produces one. Even 0.99 tops out at B+ with its realized 0.663 attached.
Raising that ceiling later is a deliberate, visible act.

Bands that cannot separate SAY so -- C+/C/C- carry
separates_from_base_rate false and copy naming it, which is the honest
description of a forecast explaining 3% of variance. Every grade states
its basis (forecast_only vs forecast_plus_matchup_factors, naming which
factors fired) and calibrated:false. engine1.grade is preserved as
engine_grade so nothing downstream breaks.

PHASE 2 — refusals render real states: insufficient_data -> "not enough
history to call this one"; juiced_no_edge -> "the book has priced the vig
past any edge on this side". 1,870 refused snapshots carry exactly those
two reasons and both now surface.

PHASE 3 — hand-verified on 12 real served props. Freeman/Rice/Encarnacion
0.95 overs now B+ (was B, C, B); the 0.05 unders now F (was C). Refused
doubles render NO READ with their reason. never-blank PASS,
no-manufactured-A PASS.

Serving change; nine frozen model modules unchanged including engine1;
p_win never mutated; no calibrated number leaks (deployed set empty); no
Bonferroni slot.

STILL TRUE: the forecast explains ~3% of outcome variance. This order did
not make the model better. It made the letter stop overstating it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 04:53:02 -04:00
builtbykev 494c83cf76 Hunt the window-bug class: three more paths, and the forward re-audit rule
in code

PHASE 0 — getStatRows is the single base-rate path, so every branch is
audited, plus the feature builders since l20_avg is the season reference
projectionFor reads:

  getStatRows MLB -> estimator base    fullLog            CORRECT (929fd81)
  mlbGameLogFeatures l5/l10/l20        last10 = 10        DEFECTIVE
  espnStatsAdapter.parseGameLog        slice(0,20)        DEFECTIVE
  getStatRows NBA/WNBA ESPN branch     inherits 20-cap    DEFECTIVE via source
  getStatRows NBA/WNBA python branch   getGameLogs(...,20) dormant (offline)
  pitcherEngine / skillProjection      statcast profiles  N/A
  pitcher props via getStatRows MLB    fullLog            CORRECT
  settleSource                         full log (S64)     CORRECT

THE PITCHER ANSWER IS GOOD NEWS: pitcher props run through the same
getStatRows MLB branch, so 929fd81 repaired them too. There is no separate
defective pitcher base-rate path.

THE ONE HIDING IN PLAIN SIGHT: mlbGameLogFeatures carries the comment
"l20 = all available (the season per-game reference projectionFor needs)"
while building from last10 -- so l20_avg was a TEN-GAME AVERAGE WEARING A
SEASON LABEL, feeding both the consistency pull inside the estimator and
projectionFor, which decides refusals. It survived the previous repair
because that fix touched only getStatRows.

PHASE 1 — mlbGameLogFeatures now reads fullLog; espnStatsAdapter drops its
slice(0,20) cap. ZERO new API calls on both: each widens data already
fetched and then discarded, the same shape as the original repair. The
python branch is left alone -- the service is offline in prod and fixing it
would be speculative.

Their before/after resolution is NOT measured, deliberately: the only way
to measure today is to reconstruct the repaired forecast over old rows,
which is the reconstruction-vs-served trap this order refuses. Code fix
now, measurement at accrual.

PHASE 2 — MODEL_VERSION bumped to engine1@2026-08-07-fullwindow, so every
forward snapshot is self-identifying (retentionService already stamps it;
no new plumbing). model/reAuditEligibility.js encodes the rule: isEligible
accepts only the repaired marker, assess counts eligible DATES not rows,
and ACCRUAL is frozen at calibration 10 / hits-lift 10 / verdict-reaudit
14 / rbi-gate 14. A test locks the invisible case -- a MIXED table of 330
rows with 30 repaired returns eligible_dates 3, not 330 rows of false
confidence. Once both generations share a table a naive count would fit a
map on a blend of two forecasters.

PHASE 3 — the board, each consequence labelled: calibration WITHDRAWN
(refits at 10 dates, never on reconstructions); factor verdicts SUSPECT
(all measured against a champion worse than a frequency table, direction
UNKNOWN, not pre-priced, 14 dates); hits factor lift UN-REMEASURABLE (10
dates, factors still wired and transmitting); rbi lineup-slot RE-QUEUED
(14 dates). Pre-registered order: calibration, hits lift, verdict
re-audit, rbi gate.

Then STOP and accrue. Nothing further can be honestly measured until the
board fills with rows the repaired champion produced.

Serving-path changes by design for the MLB feature path and NBA/WNBA logs;
eleven frozen model modules verified unchanged. p_win never mutated. No
Bonferroni slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:40:20 -04:00
builtbykev 929fd81940 Repair the champion: it was reading ten games, not a season
PHASE 0 — the defect is real past the peek. Against a FAIR point-in-time
baseline (each player's rate over games strictly before that date, >=10
prior games, box scores back to 05-01), the served champion LOSES on all
four stats, three of four CIs excluding zero:

  hits  0.00251 vs 0.00774  CI [-0.0074,-0.0011]
  TB    0.00393 vs 0.00619  CI [-0.0055,-0.0003]
  rbi   0.02481 vs 0.03133  CI [-0.0153,-0.0005]
  runs  0.00181 vs 0.00683  CI [-0.0114,+0.0008]

PHASE 1 — the cause is the WINDOW, not the weights. estimateProbability
builds its base rate as the frequency over every row it is handed, and
featureCache.getStatRows handed it res.last10. So the "season rate" was a
TEN-GAME rate, and 0.4 of the forecast was the last five OF THOSE TEN. The
0.40 recency weight costs resolution on all four stats (-0.00086,
-0.00107, -0.00562, -0.00365). Nudges are mixed and small -- harmful on
hits and rbi, marginally helpful on TB and runs -- so they are left alone.

PHASE 2 — two lines, no new data, no extra API call, because fullLog was
already fetched by the same adapter call that produced last10:
getStatRows now reads fullLog, and RECENCY_WEIGHT goes 0.40 -> 0.20.

  hits  0.00251 -> 0.00817  (tripled; now above the fair baseline)
  TB    0.00393 -> 0.00734  (above baseline; vs old CI [0.0020,0.0067])
  rbi   0.02481 -> 0.02727  (still below baseline, CI includes zero)
  runs  0.00181 -> 0.00436  (still below baseline, CI includes zero)

Gate stated exactly: hits and TB now exceed the fair baseline on the point
estimate; rbi and runs remain below but EVERY CI now includes zero, so no
stat reliably loses to a frequency table. That is a tie on rbi/runs, not a
win, and it is reported as one. Only TB's improvement over the old
champion is CI-confirmed; the rest are directional.

STALE-FIT GATE: CALIBRATION_DEPLOYED is now EMPTY. The low-param maps were
fitted on the retired forecast and fromLedger cannot rescue them -- settled
ledger rows still carry OLD p_win, so refitting today would refit the
retired forecast. Nothing is served calibrated until dates settle under
the repaired champion, and the favourite-longshot bias must be re-measured
rather than assumed to survive. The shadow duel is void.

PHASE 3 — the hits factor lift is NOT re-measured, and cannot be yet: it
needs settled rows produced BY the repaired champion, which ships in this
commit. Replaying would score the factors against a reconstruction rather
than the served forecast. Deferred, explicitly. The factors remain wired
and transmitting; only their lift is unquantified on the new baseline.

PHASE 4 — standing flag, and it is large: EVERY factor verdict in this
programme, every null and every THEATER, was measured against a champion
worse than a frequency table. Signal added to noise reads as noise. Prior
verdicts may deserve re-audit. Logged, not re-run.

Re-queued not built: rbi lineup-slot / RISP opportunity through the
two-part gate, now landing on a repaired champion.

Serving-path change by design; the byte-identical invariant inverted and
all four stats move. Nine frozen model modules verified unchanged. No
Bonferroni slot -- resolution accounting on the champion's own knobs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9
2026-08-07 03:28:33 -04:00