Review of 2026-09-25T1321Z-h021-last6-screen
Owner lane: Not recorded
What held
E1 reproduces to 1e-12 and its direction is real (10 of 12 cells better than flat, CIs excluding zero; td_oe hl2 and hurry hl2 worse). The three non-share windows (td_oe_l8, hurry_delta_l8, own_scramble_rate_l4) are nil on E2 under both row-permutation and player-clustered inference (max |t| 1.35), for both hurry definitions. Target-share and snap-share recency on targets is a real residual: three cells clear Bonf-36 under player clustering and replicate in sign in 3/3 seasons, and the record flagged it, the linear bound, and E3's blindness itself.
What did not hold
(1) 'MEASURED NOTHING end to end' for all 6 is contradicted by its own lead and by A2c: a matched-loss add on targets gains ~0.3% squared error, out of fold, control functioning. (2) 'E3_valid_candidates_significantly_better: 0' and 'F4 ... worse on all 6' are artefacts of scoring an OLS slope by absolute error on peaked, heavy-tailed residuals; they carry no information and should not have been counted toward the 24-of-35 nil tally. (3) The carry-share lead is not supported under player clustering. (4) The inference for E2 ignored player clustering, inflating the breach count from 3 to 6. (5) The hurry result is about a feature definition that the warehouse no longer holds and that an uncommitted edit is replacing; the record does not name the dependency. None of this makes the size large: the recency-contrast lead is ~0.3% of squared error on targets and nothing on absolute error or on fantasy points (E2 fantasy cluster t <= 2.12).
Finding and verdict
WEAKENED
Review scope
WEAKENED, not refuted. Upheld: nil for td_oe_l8, hurry_delta_l8 (committed definition, and current definition as sensitivity) and own_scramble_rate_l4 on the E2 correlation, and E1's direction. Not upheld: (a) 'MEASURED NOTHING' for the three season shares - target_share_season and snap_pct_season (hl2) on the targets endpoint carry a small out-of-fold squared-error gain the shipped line does not contain; (b) the carry_share_season->carries lead; (c) the E3/F4 nulls as evidence of anything; (d) the '24 of 35 nil' tally to the extent it counts E3. The gain is small (about 0.3% of targets squared error, no measurable absolute-error gain, no fantasy-point gain) and was found on rows that also selected it, so it is a lead to confirm on 2026 weeks with a matched-loss instrument, not a finding to ship. Not a claim about the 11 team windows or about the 2026 live rows. An unexplained residual of this size is information not yet found, not a floor.
Next action
(1) Any further H-021 screen uses a loss-matched instrument (OLS fit scored in squared error, or LAD scored in absolute error) and player-clustered E2; drop the MAE-scored OLS add. (2) Confirm the target-share/snap-share recency contrast on the targets endpoint on 2026 weeks 3+ (aar_forecasts vs aar_outcomes), player-clustered; carry share is not a candidate. (3) Bind the snapshot: record sha256 of the .duckdb bytes and the tempo_week/tempo_delta table hashes for any screen that reads them; the file the record names has been overwritten. (4) Commit or revert the playerweek/tempo.py edit and say which hurry_delta_l8 production serves; the stored shipped rows and every H-021 hurry number use the committed one.
Source provenance and publication scope
Owned research record: research/scientist/experiments/2026-09-26T2210Z-referee-h021-last6-screen.json
This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.