930d526b0109c9357ccbcb8386d94ee2994e6229
34 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
930d526b01 |
Lineage productization: a live graph, an observer that doesn't trust it, and a surface that owns only what it knows
The authority review found there was nothing to switch: lineage is a
write-only graph with no living permanent writer and no product consumer.
Authority theater would have been a switch on a consumer that does not exist,
reading a store that is not being written. So: make the graph live, measure it
from outside, and expose the one category it alone owns.
WRITE-PATH FAILURE SEMANTICS, TRACED FIRST. persist(:857) -> the authoritative
cacheSet snapshot:latest(:1404) -> the lineage gate(:1426) -> ledger(:1473).
Lineage failure cannot fail base retention (earlier, separate upsert), cannot
fail publication (the product write precedes the gate), cannot fail the Ledger
(lineageIndex defaults null; the ledger has its own guard), and cannot create a
new partial row (blankLineage() first, atomicity sweep after the catch). The
isolation this tranche needed already existed; only a mode was missing.
THREE MODES, ONE EVALUATOR. `lineageWriteMode` distinguishes OFF /
CANARY_LEASED / PERSISTENT_SHADOW, and the write gate and the status surface
both read it, so they cannot disagree. The canary is CONSULTED, never
converted: its <=4h absolute expiry, its dynamic evaluation and its
fail-closed parse are untouched.
AN AMBIGUOUS CONFIGURATION FAILS CLOSED. If a sport is named by both persistent
mode and an active lease, the two instructions disagree about WHEN WRITING
STOPS — the lease says 22:45, persistent says never. The dangerous reading is
the quiet one: an operator sets a bounded lease believing writing will stop
while persistent keeps it going. We cannot know which they meant, so that sport
writes nothing until the configuration says one thing. The sport allowlist is
the canary's own, so persistent mode can never widen past it.
THE OBSERVER MAY NOT ASK THE WRITER HOW IT DID. settleLedger returned
{settled:0,pending:0} — byte-identical to a healthy "nothing to settle" — while
1,444 rows sat unprocessed, and the watchdog believed it. So `lineageCoverage`
reads durable retained state only, and THE DENOMINATOR MAY NOT CONSULT
lineage_action: eligibility is "the row was published AND a natural key is
derivable from its own identity columns", neither of which the lineage path
writes. If expectation were derived from whether lineage exists, coverage would
be 100% by construction and the metric would be decoration. Zero-expected and
zero-written are kept as different answers.
THE LEGACY BOUNDARY IS OBSERVED, NOT DECLARED. `publication_id` is stamped only
by commitPublication, so the row itself says whether lineage ran. Verified on
production: 5,353 rows carry it — 4,234 complete actions plus exactly the 1,119
historical partial rows — and zero actions exist without one. No epoch constant
is invented; a date would have been a guess about when the writer was on.
publication_id NULL -> LEGACY_UNVERIFIED. Stamped but incomplete ->
LINEAGE_UNAVAILABLE, which is the honest answer for the 1,119 and is never
quietly rewritten as legacy.
THE GRADE-SHIFT BADGE IS UNTOUCHED. revised_from_grade answers "did the letter
change"; lineage answers "which published claim superseded which". Different
questions, and a test now fails if either route learns the word lineage.
Suite 397/5,496/0 · tsc 0 · 15/15 teeth.
TWO OF MY OWN TESTS WERE VACUOUS AND A TOOTH FOUND IT. Tooth 2 came back green
because the isolation tests asserted the slate was published — true whether or
not the exception propagated — while never reaching the lineage gate at all:
the fake grader never fired `onGraded`, so the collector stayed empty and
persistedRows stayed null. Fixed by firing the hook and counting the commit.
A green teeth run means the test is missing.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
f54b0627e1 |
Truthful provenance for an operator-invoked snapshot
The internal one-shot route already called the SAME production runSnapshot with
the SAME default dependencies — but it passed no trigger, and runSnapshot
defaults an absent trigger to SCHEDULED. So every operator-invoked run was
recorded as though the cron had fired it. That was a lie about provenance,
present by omission, and it would have contaminated the trace of any forced
diagnostic run.
CONTROLLED_FORCED is now its own trigger. Both internal routes (`/snapshot/:sport`
and `/snapshot/all`) stamp it, along with the process generation. The scheduler
still stamps SCHEDULED, and a test asserts neither internal route can label
itself scheduled.
Trace retention moves from "scheduled only" to a named allow-list of SCHEDULED +
CONTROLLED_FORCED. INTRADAY is still refused — it runs every ~20 minutes and
would displace scheduled evidence, which is the failure the store exists to
prevent. MANUAL_API stays refused too.
THE PIPELINE IS UNTOUCHED. snapshotService, gradeSlateService, retentionService,
snapshotScheduler, oddsService and eventIdentity are all UNCHANGED. A test
asserts the route injects no dependency override — no getOdds, gradeAndCacheSlate,
retention, ledger, cacheSet/cacheGet, gameBinder, eventIdentity, mlbAdapter or
notify — so the only difference from a scheduled invocation is the label and the
absence of a scheduled hour, which a forced run genuinely does not have.
The ?limit bisect-hook invariant is preserved and tightened: the opts passed
carry exactly {trigger, processStartedAt} and never a stray limit.
Teeth, injections verified present, against a green baseline:
:sport route mislabelled SCHEDULED -> 3 fail
/all route mislabelled SCHEDULED -> 2 fail
intraday admitted to the store -> 5 fail
trigger filter removed -> 3 fail
THE FIRST TEETH RUN WAS INVALID AND IS DISCARDED: both routes live in one file,
so a single-occurrence replace hit `/snapshot/all` and left `/snapshot/:sport`
correct — the injection landed on the wrong target and the suite passed. Coverage
for `/all` was added, plus a test that the file contains exactly two
CONTROLLED_FORCED stamps and zero SCHEDULED ones, then both were re-run failing
independently.
Four stale assertions updated with the reason recorded: three pinned the
`not_scheduled` refusal string (now trigger-agnostic) and one pinned an empty
opts object on the route.
389 suites / 5,280 tests pass. web tsc exit 0. Lineage stays OFF.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
8d741e3932 |
Pre-grading stage trace: record which branch loses the cohort
The acquisition recorder proved the previous diagnosis wrong: the 03:00 MLB
attempt acquired 6,365 props from a propline cache hit and CONTINUED. Zero
reached the first grading callback, and nothing durable says why.
WHY REPLAY WAS IMPOSSIBLE. The 03:00 input was the cache object written
02:20:24.889Z. ODDS_CACHE_TTL defaults to 3600s, so it expired ~03:20; it is
now gone. No historical copy exists: closing_captures holds no MLB rows past
01:40:52, model_snapshots none, and `bookprices:mlb` carries no timestamp the
book-comparison route exposes. Calling /api/odds/mlb was REFUSED — a cold cache
would fetch and fire recordDownstream -> gradeAndCacheSlate, writing product
state and manufacturing a false recovery. So the input is NOT RECOVERABLE and
no replay was attempted.
A SECOND CLAIM IS WITHDRAWN. The previous tranche concluded "all 6,365 props
were rejected at event admission", reasoning that dedupe cannot empty a
non-empty list. That is FALSE: `dedupeProps` also FILTERS — it drops props with
no player/stat_type/line and any prop whose book is not a MODEL book. A fully
admitted cohort can still dedupe to zero. Demonstrated in test. So the failing
branch was never established, only assumed — which is exactly what the order
forbade, and why DEDUPE_EMPTY is a first-class outcome here.
THE REGION IS SMALL AND EVERY BRANCH LOOKS IDENTICAL FROM OUTSIDE:
identity annotation (runSnapshot, mlb only, own catch)
-> admitForGrading (pure, can throw)
-> dedupeProps (pure, can throw, ALSO filters)
-> mapLimit(gradeBestSide) <- first onGraded-capable call
gradeAndCacheSlate swallows every throw in that region and returns the same
{written:false,count:0} it returns for an honest zero.
The recorder distinguishes them: READY_FOR_GRADING · ALL_REJECTED ·
IDENTITY_STAGE_ERROR · ADMISSION_STAGE_ERROR · DEDUPE_STAGE_ERROR ·
DEDUPE_EMPTY · NO_INPUT_PROPS · OTHER_PREGRADING_ERROR. An exception is never
folded into ALL_REJECTED, and with no admission evidence the classifier refuses
to classify at all — a test pins that.
EXCEPTION SEMANTICS UNCHANGED. Each stage is wrapped to record and then RETHROW
the identical error, so the enclosing best-effort catch still handles it exactly
as before: nothing caught that was not caught, nothing swallowed that was not
swallowed. Admission and dedupe rules, the impossible-binding invariant, event
identity, model books and the first-row-wins cap are untouched — the only
behavioural lines in the diff are `const gate/unique` becoming `let`.
Correlated to the SAME snapshot_attempt_id the acquisition recorder minted — not
a new run id — and stored under its own key `ops:pregrade:{sport}` so it can
never displace the acquisition record. Same atomic LPUSH/LTRIM pattern, bounded
per sport, scheduled-only, best-effort at the call site, auto-disabled under
test.
CAUGHT DURING BUILD: `savePgTrace` was defined and NEVER CALLED — the recorder
would have persisted nothing, the same "built, correct, never invoked" failure
this programme has hit before. A test now drives the real runSnapshot and
asserts a trace is persisted carrying the acquisition's attempt id.
Ten teeth, injections verified present, against a green baseline:
empty collector read as ALL_REJECTED (2) · admission throw reported as
admitted=0 (1) · dedupe throw reported as output=0 (1) · dedupe-to-zero
mislabeled as rejection (1) · identity exception hidden (1) · observer reorders
the candidate array (1) · trace failure changes product outcome (1) · intraday
overwrites scheduled (1) · different attempt id (2) · secret leak (1).
Restored byte-identically.
THREE OF THOSE LANDED AND PASSED FIRST TIME — coverage holes, not safe defects:
the dedupe-throw and identity-throw tests only exercised the recorder directly,
never the real path, and nothing asserted the CANDIDATE array is not reordered.
All three closed with real-path tests, then re-run failing.
Two stale source assertions updated with the reason recorded: both pinned the
exact `const gate = …` / opts-key order that the recording wrapper changed.
387 suites / 5,237 tests pass. web tsc exit 0. Lineage stays OFF.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
|
||
|
|
238f0f67cf |
Scheduled acquisition trace: record the fork instead of inferring it
MLB stopped producing anything at the 22:00 and 01:00 slots on 2026-08-27 while NBA/WNBA ran normally. The differential narrowed it to exactly two runSnapshot exits — getOdds THREW, or getOdds RETURNED ZERO PROPS — and production retained nothing able to tell them apart. A failed acquisition has no snapshot_id, writes no ledger row and updates no slate, so it left no durable evidence at all. The alert channel could not fill the gap either: quota/test-alert reports sent:true while the ntfy topic replays 0 messages, so absence of alerts is not evidence. This records the decisions the existing control flow already makes. It changes no acquisition behaviour: provider order, the single existing retry, the quota threshold, the fallback and cache policy are all untouched. The only behavioural line in the diff is getOdds gaining an optional recorder argument, defaulted to a frozen no-op so every existing caller is byte-identical. IT MAKES NO PROVIDER CALL. Measured: zero added fetchAllOdds/getProps/gateway calls, and the trace module contains no HTTP of any kind. Tests assert both. WHY REDIS, NOT MEMORY. server.js arms the scheduler in EVERY process and there is no lock or leader election, and a rolling deploy demonstrably serves two containers at once — so process-local evidence could be written by a container nobody later probes. Storage is LPUSH + LTRIM, which is atomic: two schedulers racing the same slot both survive instead of one silently overwriting the other, and each attempt carries a process_generation so they stay distinguishable. WHY A BOUNDED HISTORY, NOT "LAST ACQUISITION". The scheduled MLB attempt failed at the hour and an intraday attempt SUCCEEDED ~20 minutes later. A single last-value would have erased the failure with the success — precisely the evidence needed. Only SCHEDULED attempts are retained (persist refuses any other trigger), each sport keeps its own key, and intraday structurally cannot write one because intradayRefreshService never calls runSnapshot. WHAT IS CAPTURED, per attempt: cache decision; PropLine outcome as NONZERO/ZERO/ERROR/NOT_ATTEMPTED with count; the odds-api fallback with allowed_at_invocation, blocked_reason and the quota AS OBSERVED AT THAT INVOCATION — reading provider quota hours later and calling it historical evidence is the exact mistake this exists to stop; then the final result and the runSnapshot terminal outcome. The EXISTING retry appears as a second attempt; no retry was added. Sanitized: keys, tokens, URLs and long opaque strings are redacted, and no prop payload is retained — counts only. Tests assert a dirty provider error and a real prop array both come out clean. BEST-EFFORT AT THE CALL SITE, not just in the default dep — a teeth proof showed an injected store could still throw into a healthy snapshot. Now any implementation is safe. persist also auto-disables under NODE_ENV=test unless a client is injected (the opsNotify precedent); without that the default path opened a real ioredis connection inside every suite driving runSnapshot. Read-only GET /api/internal/acquisition/:sport behind the existing internal auth. It runs no pipeline and makes no fetch. Nine teeth against a green baseline, injections verified present: intraday overwrites scheduled (1) · one global slot (2) · thrown getOdds with no terminal trace (2) · zero mislabeled as error (1) · quota not captured at invocation (1) · observer makes a provider call (1) · credential leak (2) · scheduled/intraday share an identity (1) · telemetry failure breaks the snapshot (1). Restored byte-identically. TWO OF THOSE LANDED AND PASSED FIRST TIME — coverage holes, not safe defects: the provider-call scan did not forbid getOdds, and nothing exercised a non-scheduled trigger through runSnapshot. Both closed, then re-run failing. Model, retention, event identity, admission, dedupe, ledger, lineage, cadence, quota tracker and the PropLine adapter are all UNCHANGED. Lineage stays OFF. 386 suites / 5,210 tests pass. web tsc exit 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8 |
||
|
|
9809626c99 |
Retention completion: a cohort is complete only when the writer says N of N
The previous bug made the recorder write nothing. The dangerous successor is a
recorder that writes half and looks healthy: persist() writes in chunks of 250
and STOPS AT THE FIRST FAILED CHUNK, so chunks committed before the failure are
already durable. Rows exist under the snapshot_id, captured_at is uniform, Redis
kept working — and the cohort is short.
So row presence was never completion evidence, and neither was a matching
timestamp. Completeness is now proven by the writer or not at all.
TERMINAL RETENTION STATES (retentionService.classifyPersist):
NOTHING_TO_PERSIST attempted 0 — a refusal-only slate is still a cycle
SKIPPED_NO_DATABASE no database configured; not a failure
COMPLETE attempted > 0, written === attempted, no error
FAILED_ZERO_WRITE written === 0 — first chunk failed
FAILED_PARTIAL 0 < written < attempted — a later chunk failed
FAILED_UNRESOLVED_ERROR counts look complete but an error is unresolved;
unreachable through today's loop, and kept because
the alternative is reporting COMPLETE holding an error
The invariant: any written < attempted with attempted > 0 is a FAILED cycle. A
partial cohort is never degraded success.
classifyPersist reads the EXACT persist() result and refuses anything else — it
never recomputes attempted or written, because a second calculation could
disagree with the writer and then the status would describe a cycle that did not
happen. persist() itself is byte-identical to
|
||
|
|
ceaa896f77 |
Runtime observability: report the build and canary state the system acts on
The rollout stalled at RUNTIME_UNVERIFIED because two facts were answerable only
as a side effect of a scheduled snapshot writing a row: which build is running,
and whether MLB lineage is effectively enabled. Every state transition therefore
waited on cron rather than on asking the service.
- src/services/lineageCanaryConfig.js — THE canary resolver. Parsed once at
module load (unchanged semantics), normalised sorted/deduped/trimmed, frozen.
snapshotService's write gate now delegates to it, and the status probe reads
the SAME state. A route that parsed the environment itself would be a second
version of the truth, free to drift from the gate it claims to report.
- GET /api/internal/snapshot/status gains runtime.code_sha (the production
codeSha resolver — never git, never gitea/main; null when unavailable),
runtime.started_at (computed ONCE at module load, so it marks a boundary
rather than reading as now; deliberately not called deployed_at), and
lineage_canary {enabled, sports, configuration_source}.
- No raw environment value is returned; sports is the normalised set and
configuration_source says only ENVIRONMENT vs DEFAULT. Router-wide
requireInternalAuth is unchanged: 200 with key, 401 without.
- Effective lineage config is fixed for the process lifetime, so
runtime.started_at is a defensible lower bound for how long that state held.
Strictly observational — the handler still only reads Redis.
Model and decision code byte-identical to
|
||
|
|
7f69fef14c |
Add an on-demand endpoint for the lineup-context ingest
The first prod run wrote zero rows while the parser demonstrably works locally (144 rows, 10 games with lineups posted), so the zero was wiring rather than absence -- but diagnosing that required a full snapshot, which now takes about three minutes and 524s at the edge. Same reasoning as the statcast refresh endpoint: a job is proven by running it and reading the result, never by waiting for the slot it rides in. This makes the ingest verifiable in seconds, so 'zero rows' can be told apart from 'no lineups posted yet' immediately -- which is the exact confusion the defence ingest hit when a doubled path 404'd and read as 'Statcast has no fielding data'. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01W1sivYNqY2TS5ftykmHBU9 |
||
|
|
c9d56d5668 |
Audit endpoint: where the env/matchup context chain drops off
The arch-v1 environment and matchup axes fired on ZERO prod rows across 634 graded props while archetype axes fired normally, and buildContext works locally (15 games, 14 with weather, Coors composing to 1.241). So the failure is downstream of buildContext and has to be located, not inferred from an absence. Replays buildContext + contextFor against the CURRENT cached snapshot grades and counts the drop-off at each hop: team field present -> resolves to an abbr -> abbr matches a game -> environment produced; and bats / playerId / opposing-pitcher known -> matchup produced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
3e78217678 |
Step 0 input check: read-only feature-coverage probe
Before wiring any layer into the grade, measure whether its inputs are actually populated on real props. A layer wired onto sparse inputs does not degrade gracefully by default -- Number(null) === 0 turns a missing opportunity into 'zero opportunity', a fabricated input rather than an absent one. Reports population per feature, SPLIT BY stat_type, because a feature can be 100% present for batters and 0% for pitchers and a pooled number would hide exactly that. Also reports whether ab_per_game varies across a player's own props -- a per-player constant can only move all of a player's props together, which is a very different thing from a per-prop opportunity signal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
8d052131c5 |
Add a bisect hook (?limit=) to the internal snapshot trigger
The cap raise 25 -> 500 made an induced snapshot 502 at 13.4s and the run did not complete in background either, while a 25-prop run had completed in 16.3s. That rules out a simple duration timeout and means the cause has to be measured, not guessed. ?limit= bounds one run so the regression can be bisected without a prod env change; omitted, the real DEFAULT_LIMIT applies. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
d18a19f6aa |
Part 1 diagnostic: read-only refusal categoriser (25-cap + 72% refusal)
READ-ONLY. Runs the REAL grade path over a REAL slate and categorises every refusal; writes nothing. Reproduces gradeSlateService.dedupeProps exactly (MODEL_BOOKS, first-row-wins) and calls analyzeViaEngine1 the same way, so it measures what the pipeline does rather than a re-implementation. Adds a FIFTH bucket the order did not anticipate, and it is likely to change how the 72% is read: (e) POLICY-SUPPRESSION. The 2026-07-19 betting-logic audit deliberately refuses rare-event 0.5 markets (doubles/ triples/HR/SB) on the juiced under, plus any over-juiced price -- and it sets the SAME insufficient_data flag as a genuine data gap. Counting those as a data problem would send us hunting for data that is not missing, and "fixing" them would re-introduce bets we removed on purpose. Separates (b) FETCHABLE-GAP from (d) GENUINE-ABSENCE by asking the stats layer directly whether the player has ANY game log, rather than assuming: no log -> genuine absence, keep refusing; a log that exists while the grade path found no projection -> a wiring gap with something to fix. Also measures per-grade latency (mean/median/p90/max, serial and at concurrency) so Part 2 can decide the cap on cost rather than on taste. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
86d123945c |
Rank on p_win: challenger instrument + retire edge from decisions
MEASURED BASIS (n=200 settled MLB rows): corr(p_win, outcome) = +0.26; corr(edge, outcome) = -0.010 incumbent ruler / -0.022 consensus ruler. Subtracting the market destroys the signal under BOTH rulers, so a quantity that does not predict must not rank, gate or decide. CHALLENGER-FIRST -- live ordering is byte-identical. rankGrades (the incumbent, grade-first with edge as its 4th key) is untouched and tested as untouched. NEW: rankByForecast -- takeable-gated p_win -> grade -> confidence -> stable order, with NO edge term anywhere. p_win LEADS and the letter follows, deliberately: the letter measured r ~ 0.005 and is inverted (B 52.4% < C 56.9%) while p_win measures +0.26, so leading with the letter would sort by the weaker signal and use the stronger one only to break ties. Recorded in the code: isotonic calibration is a MONOTONE transform, so ranking on raw vs calibrated p_win gives the SAME ORDER. Calibration matters when p_win is displayed or thresholded; it cannot change a ranking. Nothing here needs the calibrated value. rankingDelta + GET /api/internal/ranking-delta measure how far the board would move before any flip. The endpoint reports p_win coverage alongside the delta -- if p_win is absent the challenger degrades to grade order and the delta UNDERSTATES, which is worth saying rather than reporting a clean zero. forecast_rank is stamped on snapshot grades BEFORE stripModelPrice, so every tier gets the correct order without the paid values (the topGradedService precedent -- an ordinal can travel where the magnitude cannot). Additive only: nothing sorts by it yet. RETIRED AS DECISIONS (not rankings, so done now): - altLineScanner.compareToBookImplied no longer returns value_detected: edge > 0. Edge is still COMPUTED and returned -- losing the record would be worse than mis-using it -- but the verdict is an honest null with value_basis: 'retired:edge_does_not_predict'. - scanAltLines no longer filters to edge>0 or calls the survivor "optimal". The whole ladder is returned ranked and labelled 'price_gap_diagnostic_unvalidated'. The module has ZERO callers (verified) -- unwired like mlbGrader.js, left in place and made honest. An honest asymmetry recorded there: ranking props AGAINST EACH OTHER must not use edge, but choosing between RUNGS OF THE SAME PROP is inherently price-relative -- ranking rungs by model probability alone would always pick the lowest line, since P(over 0.5) > P(over 2.5) by construction. So the gap stays the rung key, explicitly labelled unvalidated. Two superseded tests updated to stronger properties. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
2071b79456 |
Order Zero Phase 1: keyed read-only PropLine verification endpoint
Adds GET /api/internal/propline-verify (internal-key gated, read-only) so Phase 1 can run WHERE THE KEY LIVES. Touches no cache, no ledger, no grade; the live adapter and the live ruler are untouched. Breadth reuses proplineAdapter.fetchRaw -- the exact live request -- so what it measures is what the pipeline actually receives. Reports per sport (never pooled): books/prop from the feed vs after our own ALLOWED_BOOKS, props made INVISIBLE by that filter, reference-book presence, DFS presence reported separately, and consensus eligibility. Consensus eligibility is deliberately strict: >=2 REFERENCE books posting BOTH sides at the SAME line. A one-sided quote cannot be de-vigged, and two books at different lines are not the same market -- counting either would overstate how much of the slate can carry a real ruler. Probes the documented-but-unverified endpoints (/sports, /context, /odds/closing, /movement, /results, /exports/resolved-props for four sport keys) and classifies works/partial/no, with 403 = tier-gated and 200-but- empty = partial rather than works. Key safety is the other locked property: the key goes via axios params, never string-interpolated, and every emitted string passes scrubKeys() which removes the literal key AND any surviving apiKey= query value. A test asserts a thrown transport error carrying the key cannot escape. 13 unit tests, hermetic (no network, no key). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QJs13VsyiSKYQP6rj3NNmc |
||
|
|
6552281661 |
CLV instrument repair: fix attachClosingProb read + recoverable market_unavailable
The closing_prob funnel collapsed 100k priced captures -> 59 usable. Root cause (VERIFIED against prod, join key is PERFECT with 0 mismatches): - attachClosingProb read closing_captures with .limit(50000) and NO ORDER BY on a 730k-row table that is 86% refusal rows -> saw ~7% for MLB, missed most priced closes and declared 200+ rows closeless that HAD a capture. - market_unavailable_reason was write-once/terminal, so a row wrongly declared (truncated read / premature declaration before the capture was visible) could never recover even once its genuine capture existed. 298 rows (204 MLB + 94 WNBA) were stuck this way. Fix (CLV computation only — no grade/locked_odds/outcome touched): - Read ONLY priced captures (missed_reason IS NULL, both odds NOT NULL), scoped to the candidate rows' game_dates -> small AND complete, no arbitrary truncation. - Drop the market_unavailable exclusion from candidates; make it a re-checkable absence: a genuine close now UPGRADES the row (writes closing_prob, clears the verdict). closing_prob stays write-once (first true close wins). No capture + past game -> still declared absent (honest). No churn on already-absent rows. - New internal trigger POST /api/internal/ledger/attach-closing[/:sport] for backfill + verification (scheduler already runs attach per tick). Recovers ~312 usable closes (59 -> ~371), MLB included. Capture itself was healthy all along (94.9% MLB / 95.8% WNBA per-prop coverage). Full suite 3835 green (17/17 instrument tests incl. 2 new recovery cases), web build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VsztNChZ7vEvSR61AuMhD1 |
||
|
|
f33091ddb8 |
Layer 3 Step 4b: public park base, source-pluggable, plus game-level capture
PHASE 0 — the settle path sees a player's game-log line, not the game. It knows date and teams, never venue or final totals. But the grain is far cheaper than per-prop or even per-game: ONE statsapi schedule call per game DATE returns every game that day with venue, linescore and scoring plays. Fifteen games, one call, verified live. PUBLIC BASE — the ingestion was already done. The static FanGraphs table from Session 15 is the public base; this converts its 100-indexed values into the multipliers the composable architecture wants (Coors 128 becomes 1.28) rather than ingesting a second copy of a number we already hold. It is labelled COMMODITY in the code, not just in a comment. Every resolution carries a provenance record, and the public one reads proprietary: false with the note "Commodity: a public number. Not a VYNDR derivation." The proprietary label exists but belongs only to the self-derived version, and only once it beats this base on the instrument. A surface rendering a park effect can state which it is rather than implying the flattering one. Honest-absent where even the PUBLIC number is thin: a relocated club in a temporary venue gets no factor, because a public number for a park with one season behind it is no more trustworthy than ours would be. SOURCE-PLUGGABLE is the architectural point. resolveParkBase() is the only accessor, public and derived return identical shapes, and callers never branch on source — so when self-derived factors clear their floor they swap into the same slot with nothing downstream to rewrite. A derived source with no factor available returns absent rather than silently falling back to public, because a silent fallback would make a proprietary claim out of a commodity number. GAME-LEVEL CAPTURE starts now because it cannot start retroactively. Game grain, deduped on game_id, never copied onto prop rows — a game's totals belong to the game, and duplicating them per prop is how one fact starts disagreeing with itself. Every field is tied to a named future derivation: venue for park factors, runs for the run environment, HR totals for HR factors. Nothing else is stored. Only Final games are captured, since an in-progress total is not a result, and a game with no scoring plays reports HR as absent rather than zero. HR totals come from scoring plays, which is complete because every home run scores at least the batter. The accrual target is stated rather than promised: 150 home games per venue at roughly 81 per season means about two seasons before a self-derived factor can be nominated, and accrualStatus() reports live progress per venue so the wait is measurable. Induced: Coors home runs +0.061 for the hitter and identically +0.061 for the pitcher's home-runs-allowed at the same park, mirrored on the under; San Francisco negative; Tampa flagged weather-N/A with its factor still applying; the Athletics' temporary venue absent; strikeouts untouched. A real 2025-07-19 capture produced 15 games across 15 venues, 12 with HR totals, zero duplicate game ids. Migration 035. Induce with POST /api/internal/gamectx/:date, progress at /gamectx/accrual. Tests 3688 passed / 298 suites, web build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj |
||
|
|
528cb1a6d0 |
Layer 1: Statcast mechanism-data ingestion (backfill + nightly refresh)
The data foundation for the archetype and projection layers, built as the pattern every sport inherits. Layers 2 and 3 are not touched. PHASE 0 GATE — both match rates measured live, both 100%. Batters 40/40; PITCHERS 66/66 across five real rosters (CLE, DET, MIN, NYY, LAD) joined by MLBAM id against the 713-pitcher Savant feed. Zero honest-absent on identity, because the join is an integer both systems use natively — and the snapshot pipeline already stores it per graded row. SOURCE — five Baseball Savant CSV leaderboards, free and public, pulled with axios and the CSV parser savantAdapter already runs in prod. pybaseball is deliberately NOT used: it is an MIT wrapper over these same URLs, and adding it would reintroduce a Python runtime in a stack where the existing Python service is already offline. min=1 on every feed, not Savant's default min=q, so the long tail arrives and OUR minimum-sample gate decides what is thin — explicit and testable rather than silently dropped upstream. Measured: 1,354 rows per season (604 batters, 750 pitchers), all five feeds in about five seconds. Pitcher mechanism includes arm angle, GB/FB/LD, chase and whiff; batters get exit velo, launch angle, barrel and hard-hit, chase and z-swing. Handedness rides in free on the movement feed (677 pitchers); batter handedness stays absent pending a roster join rather than being guessed. BACKFILL AND REFRESH ARE THE SAME CALL — a full re-pull upserted on (sport, season, source_id). Idempotent and self-healing: a missed night self-corrects on the next run, with no incremental who-played bookkeeping to drift out of sync. At 1,354 rows the simple thing is also the robust one. HONESTY RULES, each with a test: a metric the feed did not carry is null and never 0; a thin sample is STORED and flagged rather than dropped or inflated, because thin and missing are different claims; an unjoined player is stored with a null player_key and joins later; and if every feed comes back empty the job REFUSES to write, so a bad night can never blank a good table. Freshness is treated as a truth property. updated_at on every row, and the scheduler pages on a failed run AND on silent staleness — a job that stops being scheduled never produces a failure, so staleness has to alarm on its own. Never-built is deliberately not stale: different condition, different fix, and paging on a fresh install teaches the operator to ignore the alarm. Nightly at STATCAST_HOUR_UTC (default 11 UTC, after every game is final), kill switch STATCAST=0, and induce-able at POST /api/internal/statcast/refresh with a freshness probe at /statcast/status — we verify a refresh by running it, not by waiting for the slot. Migration 030 applied. Promoted columns for the classification-critical metrics plus a metrics JSONB carrying every raw field, so Layer 2 can reach something we did not promote without a re-ingest. Raw per-pitch stays out of Postgres on purpose: one season is ~0.85 GB against a 500 MB plan ceiling, and it is re-pullable from the free source if Layer 3 ever needs it. Pattern documented in docs/MECHANISM-DATA.md for NBA tracking and NFL Next Gen. Tests 3581 passed / 292 suites, web build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VCNgGSt5qvcLxaeQqa7Zpj |
||
|
|
77a58e4113 |
Arm harness on our scheduler + start closing-line capture (capture only)
PART A — HARNESS ARMED ON OUR OWN INFRA. snapshotScheduler now runs the
nightly backtest at HARNESS_HOUR_UTC (default 14), appends to
harness_results, and pages via opsWatch.harnessStaleAlarm — a validator
that stops running looks exactly like one that keeps passing. No external
dependency: the join is plain SQL through the service client and the
harness is a pure function. POST /api/internal/harness/run induces the
same code path on demand, because a scheduled mechanism is verified by
inducing it, never by waiting for a slot.
PART B PHASE 0 — GATE PASSED for what is capturable:
- C4 diagnosed: closing_line is ONE overwritable field with no timestamp
and no provenance. captureClosing writes the current line and, when a
prop fails to match, silently leaves the earlier value (= the lock) in
place — so "captured a real close" is indistinguishable from "never
updated". It is 92% equal, not 100%: 56 rows DID record movement, so
the defect is provenance, not the value.
- Feeds: normalized props already carry BOTH raw side prices per book,
with game_time, and the intraday refresh polls every ~20 min during
slate hours — so the last observable pre-lock line is available.
- SHARP close: pinnacle is in ALLOWED_BOOKS -> a no-vig reference is
capturable ("beat the market").
- ODAWA: NOT capturable. 'odawa' exists only as a UI preference option in
onboarding/settings; it is in no adapter, no ALLOWED_BOOKS, no feed. An
un-capturable source is a finding, not a gap to paper over.
- JOIN: must drop `line` from the natural key, because a close that MOVED
off the graded line is the entire point of CLV. Verified safe — all 164
current identity groups have exactly ONE line per
(sport, player_key, stat, side, game_date). Zero ambiguity.
PART B PHASE 1 — CAPTURE ONLY, built test-first. The refusal was proven
before the capture logic existed: unbound game_time, doubleheader
ambiguity, a missed pre-lock window, or a one-sided price all record
missed_reason with NO price. A stale or mid-day line substituted for a
close would manufacture a CLV proof from a number that was never the
close.
migration 029 closing_captures: append-only, never overwritten (that is
the provenance C4 lacked), BOTH raw side prices so the existing de-vig
engine can compute a fair closing probability later, sharp vs book line
types kept distinct. Wired into the intraday refresh with a capture-rate
alarm — a missed close is unrecoverable.
NO CLV metric built, as ordered. This starts the clock.
Suite 284/3417 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
|
||
|
|
2bfae804da |
Verify off-box presence by reading the remote dir back
Exit 0 from the backup script is deliberately tied to ON-BOX durability, so it is not proof the off-box copy landed. GET /api/internal/backup/offbox runs rsync --list-only against BACKUP_REMOTE using the SAME pinned known_hosts as the push (checking never disabled) and returns the dumps actually present, with size and timestamp — so off-box presence is a verified fact rather than an inference from an exit code. Needed because the dev box cannot authenticate to the Storage Box: the authorized key installed there is Kev's ~/vyndr-backup-key, not the keypair generated in-session, so independent verification has to run from the container that does hold working credentials. Suite 280/3338 green, build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA |
||
|
|
c4c9b97604 |
Off-box backup: pin the host key, guarantee the remote dir, page on failure
PHASE 1 — HOST KEY STATICALLY PINNED. ssh-keyscan -p 23 returned an ED25519 key whose fingerprint EQUALS the out-of-band value SHA256:XqONwb1S0zuj5A1CDxpOSuD2hnAArV1A3wKY7Z3sdgM, so it is safe to pin. scripts/storagebox_known_hosts now carries that verified line and ships to the container (Dockerfile already COPYs scripts/). backup-db.sh uses StrictHostKeyChecking=yes + UserKnownHostsFile=<pin> instead of accept-new, which was trust-on-first-use and would have accepted an impostor on the very first run. A missing pin file REFUSES the push rather than silently falling back. Never weakened to accept-new/=no//dev/null — a test asserts that on executable lines. PHASE 1b — REMOTE DIR GUARANTEED. The box has only .ssh/, and rsyncing a file into a missing parent either fails or silently writes the dump AS the directory name — one file, overwritten nightly, reading as "backups exist" while retaining exactly one. Uses rsync --mkpath when available, else an explicit remote mkdir -p ahead of the push. PHASE 2b — FAILED OFF-BOX PUSH IS NOW LOUD. Off-box is required, so the failed-push path pages at "urgent" (was "low"/deferred) and the script emits a machine-readable OFFBOX_OK=1/0/deferred that POST /api/internal/backup/run surfaces as a distinct offbox_ok field. Exit code deliberately still reflects ON-BOX durability — a good on-box dump must not raise a false total-failure alarm. Surfacing the truth, not manufacturing a failure. No key material is echoed anywhere; only the PUBLIC host key is committed. Suite 280/3338 green, build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA |
||
|
|
97e4dc72d5 |
Backup: chown /app/backups in image + report uid/writability
The real backup run failed: pg_dump could not write to /app/backups — 'Permission denied'. Cause: the container runs as the non-root 'vyndr' user (Dockerfile USER vyndr) and the Coolify-mounted volume is root-owned, so the mount is present but unwritable. - Dockerfile now creates AND chowns /app/backups to vyndr alongside the existing /app/data + /app/.pm2 line. Docker seeds ownership into a NAMED volume on first creation, so this fixes it for a fresh volume; a host bind-mount still needs a host-side chown, which is why the next change exists. - GET /api/internal/backup/verify now reports process uid/gid, backup_dir_writable and the access errno, so the exact chown target is observable instead of guessed. A mounted-but-unwritable volume reads as 'configured' everywhere else — this makes it loud. Suite 278/3310 green, build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA |
||
|
|
ef7f17610f |
Backup: durable on-box volume, off-box DEFERRED, and a real read-back check
BACKUP_DIR is now a persistent volume (/app/backups), so the dump already survives redeploys — the container-ephemeral risk that made this urgent is closed. Storage Box SSH auth is not sorted yet, so the off-box push is explicitly DEFERRED rather than failing: - gated on BACKUP_OFFBOX=1 (plus BACKUP_REMOTE and BACKUP_SSH_KEY); until then the script logs "off-box push DEFERRED" and exits clean. - if an enabled push DOES fail, it is a LOW-priority "deferred" notice, not a failure — the durable on-box dump succeeded, and calling that an incident would train us to ignore backup alerts. Adds the read-back check, because a backup nobody has read is a hope: countRowsInDump() runs `pg_restore --data-only --table=X -f -` and counts the rows between `FROM stdin;` and the terminating `\.`, proving the archive CONTAINS the data rather than merely parsing. Needs no Postgres server, so it runs inside the API container. Validated against a real pg_dump from a scratch Postgres: counted exactly 604 rows. GET /api/internal/backup/verify exposes it (newest dump in BACKUP_DIR, size, table, rows_in_dump). Unit tests inject spawn/fs so CI needs neither docker nor pg_restore. Suite 278/3310 green, build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA |
||
|
|
40aba37f83 |
Backup: env-injected SSH key, nightly off-box push, triggerable run
Closing the backup for real. Three changes, each fixing something that
would have made the Storage Box target fail or silently rot.
1. SSH KEY COMES FROM ENV, not from the container. Generating a keypair
inside the API container was the obvious move and it is wrong: the
container filesystem is ephemeral, so the key dies on the next
redeploy and the off-box push starts failing silently. backup-db.sh
now reads BACKUP_SSH_KEY (a Coolify secret), writes it to a 0600 temp
file per run, and removes it on exit via trap.
2. PORT 23, verified live. Hetzner Storage Box runs full OpenSSH on 23;
port 22 answers with mod_sftp (SFTP only). Banner-checked both against
u635423.your-storagebox.de. rsync now uses
-e "ssh -p ${BACKUP_SSH_PORT:-23} ... -i <key>"; the old invocation had
no -e at all and would have gone to 22.
3. OFF-BOX PUSH IS NIGHTLY, not Sundays-only. A weekly push meant up to
six days of dumps existed ONLY inside an ephemeral container, which is
the same as not existing. Alert copy updated to say exactly that when
the push fails or is skipped.
Also adds POST /api/internal/backup/run (internal-key gated) so a real
backup can be TRIGGERED and OBSERVED — it returns exit code, duration,
output tail, and whether the remote + ssh key are configured. The backup
can only run where SUPABASE_DB_URL and the Supabase route live (this
container), and there was no way to fire or inspect it without a shell.
Connectivity established this session: Storage Box reachable from the dev
box on 22/23; Supabase :5432 NOT reachable from WSL2 (so the dump must
run in-container, as designed); docker IS available locally, so the
restore-verify can run against a scratch Postgres using the real dump.
Suite 278/3305 green, build exit 0.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SmNjJAwEnqHPtXbvSZR8kA
|
||
|
|
2d413cfe1e |
Quota guard: close the silent odds-api drain + reserve floor for MLB
Diagnosis (why 500/500 went unpaged): the only regular odds-api burner was
futuresService, which called axios DIRECTLY — bypassing the gateway, so it
never hit recordCall (the ONE place the WARN/BLOCK pager fires) and never
respected the 95% block. It only syncFromHeaders, which updated the counter's
number SILENTLY. oddsService (which does go through the gateway) only touches
odds-api when PropLine fails, so recordCall for odds-api effectively never ran.
Result: the counter could reach 100% with neither pager firing.
Fixes (a silent drain is now impossible, not just guarded):
- futuresService routes through gateway.fetch('odds-api', …) → counted, blocked
at 95%, and reserve-gated. Closes the raw-axios bypass.
- Reserve floor in the gateway: a DISCRETIONARY call (futures/soccer) passes
reserve=ODDS_API_RESERVE (default 50) and is refused while remaining <= reserve.
The ESSENTIAL MLB prop-backup passes no reserve and may spend to the 95% block.
→ a futures/soccer drain can NEVER starve MLB's backup path.
- quotaTracker.syncFromHeaders (the AUTHORITATIVE number) now fires the same
once-per-period WARN/BLOCK alert on a crossing — extracted fireThresholdAlert
shared with recordCall. The header-only drain now pages.
- POST /api/internal/quota/test-alert (internal-key) test-fires the pager
end-to-end so ntfy delivery is verifiable on demand.
Also (reality-corrected cadence): WNBA restored to the full grid. 2026-07-15
had two AFTERNOON WNBA games finished before the 22 UTC slot — 14 UTC (10am ET)
is the only slot early enough for a 1pm ET game's props, and on PropLine the
extra slots cost a rounding error. Soccer stays the only trimmed sport (the
real odds-api discipline). Assumption corrected by observed data.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
fc0c7bfd18 |
Merge S7 (a1): newsletter — THE VYNDR REPORT
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> # Conflicts: # BUILD-STATE.md # CLAUDE.md |
||
|
|
0e7871ac0e |
S4 (a1): the media engine — VOICE templates, /desk, Ghost drafts
4a VOICE v1.1 committed (board start); lint is EXECUTABLE — banned list + no-exclamation law enforced in the engine (throws in test, drops in prod) and locked by tests. Curly-apostrophe variants covered. 4b mediaEngine: deterministic templates (MORNING WIRE, SIGNAL, STREAK WATCH, THE SETTLE, RECEIPTS, ARCHETYPE WATCH, LINE DISPATCH) filled ONLY from snapshot/ledger/streaks JSON. Record percentages never render under n>=20 (counts + 'Record building' below). Stark layer = curated committed library (content/stark-lines.json), day-rotated selection — selected, never generated. 4c /desk (founder-only: requireAuth + DESK_OWNERS email allowlist, deny-by-default): all formats as text + <=280-char pre-segmented tweets with per-tweet copy buttons + char counts, wire/numbers-only variants, DATA BRIEF block (structured day numbers) with copy-for-claude.ai. ntfy ping after the day's first snapshot: 'Desk pack ready'. 4d ghostPublisher: DRAFTS ONLY (status:'draft' test-locked), env-gated no-op, HS256 JWT via node crypto (zero new deps). POST /api/internal/ghost/drafts saves slate preview + settle drafts. Nothing anywhere auto-posts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
32d7200571 |
S7 (a1): newsletter — THE VYNDR REPORT
Email capture + daily report assembly + operator-triggered Listmonk send.
- NewsletterCapture (dark terminal, mono) on landing (below FAQ) + /welcome
(the real signup success surface); double-opt-in note; 'Signups open soon'
when Listmonk env is unset.
- POST /api/newsletter/subscribe: public, 10/min IP limit, honeypot,
server-side email validation, forwards to Listmonk subscribers API with
preconfirm_subscriptions:false (Listmonk sends the confirmation).
No env -> calm 200 { ok:false, reason:'not configured' }. Next proxy
web/src/app/api/newsletter/subscribe/route.ts (S25 rule).
- newsletterService.buildDailyReport: signals from snapshot:{sport}:latest,
STREAK WATCH via rosterLogs -> streaksService -> streakLens, THE RECORD via
ledgerService.getModelAggregate (percentage only when hit_pct != null —
n>=20 gate — else 'RECORD BUILDING · N pending'). RG footer (21+,
1-800-GAMBLER, Listmonk-native {{ UnsubscribeURL }}) in html + text.
VOICE v1.1 lint locked by tests: no '!', no banned vocabulary, numbers
only from injected pipeline data.
- sendDailyReport: creates + starts a Listmonk campaign; env-gated no-op;
refuses an empty report. Deliberately UNSCHEDULED — only
POST /api/internal/newsletter/send (internal key) triggers it.
- docs/NEWSLETTER.md: box-side Listmonk runbook (install, double-opt-in
list, API user, Coolify env, test-send).
- Spec: specs/feature-a1-s7-newsletter.md.
Tests 2398 -> 2429 (207 suites, all green); next build exit 0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9ae64516e6 |
Session D (night2): Phase 2.5 — intraday refresh + directional movement
Odds-only refresh every 20 min during slate hours (noon–midnight ET), in-process (INTRADAY_REFRESH=0 kill switch; skips full-snapshot slots): - signed delta RELATIVE TO THE GRADED SIDE; direction is the signal. - WITH the grade → STEAM badge, never a re-grade. - AGAINST >=1.0 → re-grade THAT PROP ONLY at the real current line: holds/refuses → VALUE (better entry, same read); drops → PUBLIC revision (grade updates, revised_from_grade preserves the ORIGINAL forever, UI strikethrough on slate strips + ledger cards). - Every run recaptures closing_line/odds → the close is now refresh- fidelity; SYNC goes live by dropping SNAPSHOT_EXPECTED_INTERVAL to 1200. - Ticker MOVE events feed from the refresh (>=1.0 moves). - QUOTA MATH: 36 runs/day/sport x 4 sports <= 144 PropLine calls/day vs 9,000/day free capacity (3 keys x 3,000). Re-grades are internal compute. - POST /api/internal/refresh/all for manual runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d296e40cb6 |
Session 58: Phase 1 — Truth Infrastructure (2327 tests)
ledger_entries is live (migration 019 applied to prod, RLS + NULLS NOT
DISTINCT dedupe verified against the real database). Every grade now
persists, settles against the real result, and carries closing-line value.
- ledgerService: pipeline pre-grade upserts (public model record, user_id
null, idempotent), closing capture on every snapshot (last write before
game start = the close), settlement with SIGNED CLV (over = locked -
closing; beat/faded/flat), 30d model aggregate with the hard n>=20 rule.
- Write paths: snapshotService -> ledger (priority path); Next /api/scan ->
ledger for authenticated users only (anon never touches the public
record). Refused reads write nothing and don't burn a scan.
- Honest refusal (work-order 1.5): no projection => insufficient_data,
grade null, "INSUFFICIENT DATA - no read" UI. The web gradeAdapter no
longer displays the line as the model projection (the audit's
model==line / +0% edge degenerate); the card renders absent states.
projectionFor is sport-aware (l5 -> l20 -> {stat}_per_90 -> xG).
- /ledger: MY READS | MODEL tabs; model header shows hit% + beat-close%
only at n>=20, else RECORD BUILDING + live pending count. ModelRecord
deferred-render strip on landing + player hero. CLV + outcome chips,
revised_from_grade strikethrough (Phase 2.5 ready).
- SYNC (Task 5): thresholds vs SNAPSHOT_EXPECTED_INTERVAL (normal <1.5x,
amber >=1.5x, STALE red >=3x) via /api/snapshot/summary.
- Phase 2.5 logged in specs/vyndr-roadmap.md (build after Phase 3).
- Data-semantics hardening: strict null-safe numeric parsing everywhere a
market value is handled (Number(null)===0 would have fabricated lines).
Backend 2309 -> 2327 tests (201 suites), web build exit 0.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
2ae8a5697e |
Session 56: Full audit — PropLine + boxscore + pipeline + sport coverage (2289 tests)
Research (verified against live MLB Stats / ESPN / The Odds APIs): - specs/propline-audit.md — every stat_type mapped against our 4-layer pipeline; real MLB boxscore fields; sport coverage status; pipeline gap analysis. - specs/vyndr-roadmap.md — priority-ordered Sessions 57–64 + coverage targets. - scripts/propline-audit.js + specs/audit-data/ (raw capture). Headline bug: oddsNormalizer mapped batter_rbis → 'rbis' while the whole grade/feature/outcome chain keys on 'rbi' — every PropLine RBI prop silently failed to grade AND settle. Fixed (+ regression test). Phase 4 — wired missing MLB stats end-to-end: - PropLine MLB markets 6 → 12 (+runs, walks, doubles, earned_runs, hits_allowed, outs — same request, no extra quota). - doubles/outs/triples added to featureCache + outcomeService MLB_LOG_FIELD and all three grade whitelists (analyze/scan/validation.py). Phase 6 — pipeline resilience: - opsNotify.js: ntfy alerts (never throws, test-disabled). Snapshot success/ stale/failure alerts; retry-once on hard odds error (not on empty slate). - Missed-cron watchdog (mostRecentExpectedSlot/isSnapshotOverdue); status probe now returns `overdue`. Coverage truth: MLB is the only end-to-end-live sport; outcome settlement is MLB-only (WNBA/NBA/soccer never settle) — documented as the #1 roadmap gap. Backend 2276 → 2289 tests (+13). Web build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d09a06c054 |
Session 55: Self-learning loop + real-time layer (2274 tests)
Product overhaul core — the two transformative, differentiated systems: Self-learning loop (Phase 2): outcomeService settles locked snapshot grades against real MLB Stats API results → hit/miss/push, rolling accuracy by grade tier (30d window). Idempotent, injectable, unit-tested. New GET /api/accuracy + /api/ledger/accuracy + internal settle triggers + cron hook. AccuracyBadge (dashboard/scan/landing) is honest — "LEARNING" below MIN_SAMPLE, never a fake number. Settled HIT/MISS chips overlay the live slate. Real-time layer (Phase 1): Slate silent 60s auto-refresh (no flash, no wipe on transient blips) + "SIGNAL LIVE · UPDATED Xs ago" freshness strip; Ticker LIVE badge that flashes on fresh events. Landing (Phase 3): TopSignals shows tonight's real top-3 A-rated grades + live accuracy — the product shown, not described. Founder pricing: FOUNDER_CODE_EXPIRY default 2026-06-30 → 2026-12-31 (had lapsed, disabling every founder code + the ClaimMeter pitch). That expiry — not a tier change — was the real cause of the 4 stripe test failures. Backend 2255 (4 failing) → 2274 (all green; +19 new, +4 fixed). Web build exit 0. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
cdedecf55b |
Session 52: Coming Soon teaser + infrastructure verification (2239 tests)
Phase 1 — Push-to-Book teaser (feature not live; teaser only): - StatStrip: "BOOK IT ⟶" per graded prop (hover: "Push-to-Book coming soon"). - GradeResultCard: "PUSH-TO-BOOK · COMING SOON" footer. Phase 2 — infrastructure verification: - snapshotScheduler logs armed AND disarmed state (incl SNAPSHOT_CRON) so container logs disambiguate off-vs-crashed. - NEW GET /api/internal/snapshot/status (internal-key gated): cron_armed, cron_hours_utc, last_snapshot per sport (gradeCount/deltaCount), redis_keys existence map, ticker_count. The post-deploy pipeline health probe. - Finding: Redis AOF/RDB persistence is a server-side (Coolify) config the app can't set/verify — documented. Phase 3 — delta pipeline (verified sound, no fix needed): - runSnapshot already rotates :latest->:previous and diffs locked lines; added opt-in SNAPSHOT_DEBUG=1 [deltas] log + a trace test asserting :previous is preserved verbatim and the delta math is correct. Backend 2234 -> 2239 tests (+5), 192 suites. Web build clean (exit 0). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
f8b120c0aa |
Session 45: Snapshot pipeline + GameCard swap + live ticker (2100 tests)
The on-demand "Read" grade model is RETIRED. A scheduled pipeline pre-grades the
slate, locks grades to the line, tracks movement; the dashboard shows them already
there. Orchestrates existing services — nothing rebuilt.
- snapshotService.runSnapshot(sport): getOdds → gradeAndCacheSlate → classify
archetype per player → lock gradedAt → line deltas vs previous snapshot → write
snapshot:{sport}:latest/previous + grades:{sport} → ticker events. Fully
injectable, zero-network unit tests. runAllSnapshots = cron entrypoint.
- Internal trigger POST /api/internal/snapshot/:sport + /all (requireInternalAuth).
In-process cron (SNAPSHOT_CRON=1, UTC 14,19,22,1,3) in server.js, no new dep.
- Public reads: GET /api/snapshot/:sport (cache-only) + GET /api/ticker (merges
TICKER_MANUAL pins) + Next proxies.
- GameCard swap: live Slate renders vyndr/GameCard (legacy kept for types only),
overlays locked grades onto game props → player name once + archetype badge +
"Graded Xh ago at -115 · Current 2.5 · ▲ TOWARD +1.0". Ungraded → "Awaiting next
scan", NO Read button. On-demand onGrade flow deleted.
- Ticker polls /api/ticker every 30s, graceful fallback to hardcoded items.
- NBA/WNBA: espnStatsAdapter free fallback (defensive parse → found:false on shape
mismatch) wired into resolvePlayerStats after the offline Python service.
Env: PROPLINE_API_KEY_1/2/3, VYNDR_INTERNAL_KEY, SNAPSHOT_CRON=1, TICKER_MANUAL.
Backend 2061 -> 2100 tests (+39), 173 suites. Web build clean (exit 0).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
9b10bb4138 | Session 20: Provider intelligence — quota tracker, gateway with fallback cascade, admin quota dashboard (1476 tests) | ||
|
|
0e3839a90a | Session 18: Admin dashboard + Tank01 prefetch endpoint (1443 tests) |