27 Sep 2026 · 10:06 AM MTUpdated…
Authored research record

Research output

← All research outputs

refereeReviewed

Review of 2026-09-23T1636Z-h067-coach-zero-price

Owner lane: Not recorded

Reproduction

REFUSED. The source record has input_artifacts [] and no registration_receipt, author_id or attempts; the as-run snapshot sci-h067.duckdb is gone from pw-snapshots; playerweek/arms.py is e782a10b now vs d654b285 bound (early_season default argument, inert for rows_for/score_arm); statline.py moved (144 lines) after 17fe17b7; rel_coach was rewritten (ingested_at 2026-09-23T17:19Z, after the 16:41Z run); shipped MAE differs in the 4th decimal (3.7204 vs 3.7201, RB 4.0258 vs 4.0244) with NO playerweek/ code change between e26341a8 and HEAD, so the drift is data. Bound and unchanged: prereg 4a159f30 (mtime 10:37:20 MT, 12s after the runner started and before it finished; outcome blindness is the record's own attestation), runner 39e29596, source out.json c20a6421, source record 8fff7b5d, playermodel.py eeeb6a3c and share.py 37758ac8 byte-identical to the recorded hashes.

Finding and verdict

WEAKENED

Review scope

Headline survives, three secondary claims do not. SURVIVES (non-exact reproduction at HEAD d8c149dc on a fresh snapshot): zero - shipped pooled +0.0137 [99% +0.0023, +0.0257] (record +0.0156 [+0.0038, +0.0278]), interval excludes zero, point above the frozen +0.005 refutation line; RB +0.0303 [+0.0046, +0.0578] is again the only position that excludes zero; replica max|diff| 0.0; n 20,523 and lag coverage 5765/5349/5394 identical to the record. DO NOT SURVIVE: (1) 'carry-forward recovers all but +0.003': the record never tested zero - lag, and the interval it cites (lag - shipped spanning zero) is not an equivalence. Reproduced lag - shipped +0.0045 [-0.0025, +0.0117] (upper edge is 85% of the zero cost); zero - lag +0.0092 [-0.0016, +0.0211] at 99% and [+0.0009, +0.0182] at 95%; lag's recovered share of the zero cost is 67% with a 95% bootstrap interval [14%, 108%]. 'Cheapest writer is a carry-forward' is a recommendation the data leave open, not a measured recovery. (2) 'about 40% of the zero arm's cost is what the coach information is worth, the rest is out-of-distribution 0': not reproduced (shuffle - shipped +0.0023 [-0.0076, +0.0120] is 17% of the zero cost here vs 40% in the record, both intervals span zero), and a serve-time swap of a feature's values does not price the information in the feature in any case (that needs a refit without it). (3) 'behaves like a regime label for an unmatched/novel sideline' is asserted with no intervention; the record's own signed-move prediction failed and the replacement story was not tested against the rival that trees clamp an out-of-range 0 to the lowest bin.

Next action

Before any carry-forward is built: (a) test zero - lag on the first post-flip board, or on more seasons, with an interval that can separate them (this sample's 99% half-width is ~0.012 against a 0.009 gap); (b) run the record's NOT-RUN arm, a refit without coach_pos_* on the same walk-forward, which is the only design that prices the coach information; (c) bind how rel_coach(c,S,w) is built and that it excludes week w.

Source provenance and publication scope

Owned research record: research/scientist/experiments/2026-09-26T0800Z-referee-h067-coach-zero-price.json

This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.