2 Commits

Author SHA1 Message Date
builtbykev f54b0627e1 Truthful provenance for an operator-invoked snapshot
The internal one-shot route already called the SAME production runSnapshot with
the SAME default dependencies — but it passed no trigger, and runSnapshot
defaults an absent trigger to SCHEDULED. So every operator-invoked run was
recorded as though the cron had fired it. That was a lie about provenance,
present by omission, and it would have contaminated the trace of any forced
diagnostic run.

CONTROLLED_FORCED is now its own trigger. Both internal routes (`/snapshot/:sport`
and `/snapshot/all`) stamp it, along with the process generation. The scheduler
still stamps SCHEDULED, and a test asserts neither internal route can label
itself scheduled.

Trace retention moves from "scheduled only" to a named allow-list of SCHEDULED +
CONTROLLED_FORCED. INTRADAY is still refused — it runs every ~20 minutes and
would displace scheduled evidence, which is the failure the store exists to
prevent. MANUAL_API stays refused too.

THE PIPELINE IS UNTOUCHED. snapshotService, gradeSlateService, retentionService,
snapshotScheduler, oddsService and eventIdentity are all UNCHANGED. A test
asserts the route injects no dependency override — no getOdds, gradeAndCacheSlate,
retention, ledger, cacheSet/cacheGet, gameBinder, eventIdentity, mlbAdapter or
notify — so the only difference from a scheduled invocation is the label and the
absence of a scheduled hour, which a forced run genuinely does not have.

The ?limit bisect-hook invariant is preserved and tightened: the opts passed
carry exactly {trigger, processStartedAt} and never a stray limit.

Teeth, injections verified present, against a green baseline:
  :sport route mislabelled SCHEDULED -> 3 fail
  /all route mislabelled SCHEDULED   -> 2 fail
  intraday admitted to the store     -> 5 fail
  trigger filter removed             -> 3 fail

THE FIRST TEETH RUN WAS INVALID AND IS DISCARDED: both routes live in one file,
so a single-occurrence replace hit `/snapshot/all` and left `/snapshot/:sport`
correct — the injection landed on the wrong target and the suite passed. Coverage
for `/all` was added, plus a test that the file contains exactly two
CONTROLLED_FORCED stamps and zero SCHEDULED ones, then both were re-run failing
independently.

Four stale assertions updated with the reason recorded: three pinned the
`not_scheduled` refusal string (now trigger-agnostic) and one pinned an empty
opts object on the route.

389 suites / 5,280 tests pass. web tsc exit 0. Lineage stays OFF.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-28 02:56:22 -04:00
builtbykev 238f0f67cf Scheduled acquisition trace: record the fork instead of inferring it
MLB stopped producing anything at the 22:00 and 01:00 slots on 2026-08-27 while
NBA/WNBA ran normally. The differential narrowed it to exactly two runSnapshot
exits — getOdds THREW, or getOdds RETURNED ZERO PROPS — and production retained
nothing able to tell them apart. A failed acquisition has no snapshot_id, writes
no ledger row and updates no slate, so it left no durable evidence at all. The
alert channel could not fill the gap either: quota/test-alert reports sent:true
while the ntfy topic replays 0 messages, so absence of alerts is not evidence.

This records the decisions the existing control flow already makes. It changes
no acquisition behaviour: provider order, the single existing retry, the quota
threshold, the fallback and cache policy are all untouched. The only behavioural
line in the diff is getOdds gaining an optional recorder argument, defaulted to
a frozen no-op so every existing caller is byte-identical.

IT MAKES NO PROVIDER CALL. Measured: zero added fetchAllOdds/getProps/gateway
calls, and the trace module contains no HTTP of any kind. Tests assert both.

WHY REDIS, NOT MEMORY. server.js arms the scheduler in EVERY process and there
is no lock or leader election, and a rolling deploy demonstrably serves two
containers at once — so process-local evidence could be written by a container
nobody later probes. Storage is LPUSH + LTRIM, which is atomic: two schedulers
racing the same slot both survive instead of one silently overwriting the other,
and each attempt carries a process_generation so they stay distinguishable.

WHY A BOUNDED HISTORY, NOT "LAST ACQUISITION". The scheduled MLB attempt failed
at the hour and an intraday attempt SUCCEEDED ~20 minutes later. A single
last-value would have erased the failure with the success — precisely the
evidence needed. Only SCHEDULED attempts are retained (persist refuses any other
trigger), each sport keeps its own key, and intraday structurally cannot write
one because intradayRefreshService never calls runSnapshot.

WHAT IS CAPTURED, per attempt: cache decision; PropLine outcome as
NONZERO/ZERO/ERROR/NOT_ATTEMPTED with count; the odds-api fallback with
allowed_at_invocation, blocked_reason and the quota AS OBSERVED AT THAT
INVOCATION — reading provider quota hours later and calling it historical
evidence is the exact mistake this exists to stop; then the final result and the
runSnapshot terminal outcome. The EXISTING retry appears as a second attempt; no
retry was added.

Sanitized: keys, tokens, URLs and long opaque strings are redacted, and no prop
payload is retained — counts only. Tests assert a dirty provider error and a
real prop array both come out clean.

BEST-EFFORT AT THE CALL SITE, not just in the default dep — a teeth proof showed
an injected store could still throw into a healthy snapshot. Now any
implementation is safe. persist also auto-disables under NODE_ENV=test unless a
client is injected (the opsNotify precedent); without that the default path
opened a real ioredis connection inside every suite driving runSnapshot.

Read-only GET /api/internal/acquisition/:sport behind the existing internal
auth. It runs no pipeline and makes no fetch.

Nine teeth against a green baseline, injections verified present:
intraday overwrites scheduled (1) · one global slot (2) · thrown getOdds with no
terminal trace (2) · zero mislabeled as error (1) · quota not captured at
invocation (1) · observer makes a provider call (1) · credential leak (2) ·
scheduled/intraday share an identity (1) · telemetry failure breaks the snapshot
(1). Restored byte-identically.

TWO OF THOSE LANDED AND PASSED FIRST TIME — coverage holes, not safe defects:
the provider-call scan did not forbid getOdds, and nothing exercised a
non-scheduled trigger through runSnapshot. Both closed, then re-run failing.

Model, retention, event identity, admission, dedupe, ledger, lineage, cadence,
quota tracker and the PropLine adapter are all UNCHANGED. Lineage stays OFF.

386 suites / 5,210 tests pass. web tsc exit 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CQJeAG8vcDoL5zkiaJyVb8
2026-08-27 22:29:42 -04:00