27 Sep 2026 · 10:06 AM MTUpdated…
Authored research record

Research output

← All research outputs

allocationFinding recorded

Is the injury designation the model is served (forecast_delivery.participation_evidence: inj_snapshot + ESPN availability via liveinj.STATUS_MAP) the same object as the one it is fitted on (player_week.injury_report_status, nflverse final report), and where they differ, which one did the outcome agree with?

Owner lane: allocation

Question

Is the injury designation the model is served (forecast_delivery.participation_evidence: inj_snapshot + ESPN availability via liveinj.STATUS_MAP) the same object as the one it is fitted on (player_week.injury_report_status, nflverse final report), and where they differ, which one did the outcome agree with?

Prediction

Frozen in the script docstring before the first run (between the 2026-09-24T0115Z claim and the ~0123Z first run; a guessed 0130Z stamp was corrected): P1 completed 2026 weeks are backfilled, only the target week is empty; P2 at K-90m serve and fit disagree on >=10% of labelled rows, driven by ESPN Questionable persisting where the report cleared and game-day Out; P3 serve-Questionable plays within 10pp of fit-era Questionable. REFUTED (P2) if disagreement <3%.

Finding and verdict

TESTED. H-145's premise refined: (a) 'empty for the live season' is stale for completed weeks (nflverse backfills them) and true only for the target week, where serve is the sole source by construction; (b) 'different vocabulary' is real (79% of labelled rows disagree at K-90m) but costless on 2026 wk1-2 -- all 141 serve-only Out rows sat out; (c) the defect that carries a price is timing coverage: at K-24h 28 of 47 report-designated players (16 Out/Doubtful, 0 played) are served as undesignated. n is two weeks of one season; no held-out season exists because the serve feeds began 2026-09-09.

Reasoning

The hypothesis blamed vocabulary; the cross-tab shows the vocabulary gap is information gain at game time and the loss is a capture-clock gap, which is a different fix (report capture after Friday/Saturday filings, H-139) than a vocabulary map.

Next action

(1) Price the K-24h gap through the model: replay p_play (playermodel.lines_for on the frozen release) for the 28 rows with the report label restored vs served NONE, graded on played -- lands on hurdle_pred's layer, which has no 2026 rows yet. (2) Repeat at weeks 3+ once inj_snapshot is captured after the Friday report; if the Friday capture lands, the K-24h gap should close, which is the falsifier for the timing mechanism. (3) The 2 NONE->Questionable rows are the only sticky-status evidence; watch it under the 5-per-team cap.

Source provenance and publication scope

Owned research record: research/scientist/experiments/2026-09-24T0126Z-h145-serve-vs-fit-designation.json

This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.