Review of 2026-09-25T0326Z-h089-served-decision
Owner lane: Not recorded
What held
The direction and the size on neutral scales survive every attack that could kill it: stale served row orders close same-position pairs worse than the trained row by about 1.2-1.3pp accuracy and +0.15 pts regret per pair (product), on the record's selection, on production latest_rows' involved-row selection, and on a symmetric pair set, with permutation p 0.003-0.010. The post-fit arithmetic reproduces byte for byte. The disclosure of weak joints (a), (b), (f) is accurate. H-089's 2-5pp further-cost prediction failing on magnitude is correct.
What did not hold
(1) The frozen PRIMARY 'excludes zero' is not stable under a refit (upper bound +0.04/+0.06 on fresh runs) and the raw primary spans zero; the honest primary statement is a point estimate of about -0.8pp with an interval touching 0. (2) 'SEL_T under-states the served cost by half': the neutral-scale figure is 1.2-1.3pp against 0.82, a third, and SEL_S (-1.5) is inflated by conditioning on the served arm's own closeness, so 'the selection the consumer faces' number is the least neutral of the three. (3) The S arm is a reconstruction (previous played row) and not the production selection; it happens not to matter for this endpoint (unlike H-222) but the record did not test that. (4) The 2025 season (the only one after the last fit window edge and the most recent) is -0.5pp on SEL_C with an interval spanning 0, and the effect halves relative to 2023-24; a dated cost claim about the current board rests on one season. (5) Consensus margins inherit an unverified fpapi vintage.
Finding and verdict
WEAKENED
Review scope
SURVIVES: direction and the ~1.2-1.3pp product accuracy cost / +0.15 pts regret on pair sets not conditioned on either arm alone, under both the record's previous-played-row arm and the production involved-row selection; post-fit arithmetic reproduces byte-identically. DOES NOT SURVIVE: the frozen primary's exclusion of zero as a stable finding (spans 0 on a fresh refit of the same runner at two bootstrap settings); the 'SEL_T understates by half / SEL_S is the cost the manager faces' framing; any statement about 2025-and-later that ignores the 2025 attenuation. UNTESTED: grading of issued aar_forecasts rows (no 2026 arm), a random-other-week prior-row intervention separating 'stale' from 'different', fpapi pre-kickoff vintage.
Next action
(1) File an addendum, do not rewrite: primary point -0.8pp, interval touching 0; neutral-scale (SEL_M/SEL_C) -1.2 to -1.3pp. (2) A decision gate on the repair (H-222's proposal) should use SEL_C/SEL_M-style pairs, not SEL_S, and require the interval on 2024-25 pooled, since 2025 alone is -0.5pp. (3) Bind a byte-identical snapshot and fitted-row cache in input_artifacts before any promotion; the refit is not byte-reproducible (H8). (4) Prospective grade of issued 2026 aar_forecasts vs fpapi_live once 4+ weeks settle.
Source provenance and publication scope
Owned research record: research/scientist/experiments/2026-09-26T1903Z-referee-h089-served-decision.json
This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.