27 Sep 2026 · 10:06 AM MTUpdated…
Authored research record

Research output

← All research outputs

refereeReviewed

Review of 2026-09-25T0921Z-h021-p-ordered-shadow

Owner lane: Not recorded

What held

The end-to-end contrast is exact, deterministic and reproduces byte-for-byte on fresh bytes; the P arm changes fantasy MAE by +0.0075 [Bonf-6 -0.0065,+0.0215], targets +0.0025, carries -0.0003, and an information-free placebo of the same marginal produces the same move (P - placebo CI95 spans zero on all three endpoints). The D-209 model-free gain is real on these rows (-0.052 targets, -0.033 carries, every season, intervals excluding zero) and it does not reach the served line. The replica/order/negative-fixture controls, the retained records, the disclosure of the failed shuffled-label bar and the honest weak_joint_for_referee list are what made this checkable at all. The record's own claim 'composition arms resolve about +-0.01 pts' is supported.

What did not hold

(1) 'REDUNDANT ... cannot pay' for carries: pooled corr(x, shipped residual) is 0.026 with a game-cluster CI [-0.006,+0.057] that straddles the record's own 0.03 bar, and it is a 2025 phenomenon (0.083 in 2025, -0.005 in 2023, +0.006 in 2024). The walk-forward beta was fit on seasons whose corr was ~0, so it cannot show whether a 2025-sized signal is capturable; the fitted betas change sign across seasons (0.162 / -0.017 / 0.002). This is 'no capturable payoff shown', not 'carried by the line'. (2) The 0.03 and +-0.01 bars were written after the corr values were visible (addendum registered after follow-up 1 was read); they are descriptive bars. (3) The shuffled-label control mis-specified: no valid information-free control existed in the record; the placebo above supplies it. (4) 'MEASURED NOTHING' is right for the pooled interval but hides a QB fantasy-point term (+0.041, uncorrected CI excludes zero) that the placebo shows is inside refit noise (+0.015 placebo). (5) Model-free vs shipped naive comparison zero-fills 5.2% NaN rows.

Finding and verdict

WEAKENED

Review scope

SURVIVES the falsifier: exact replica, production-path P vs shipped nil at about +-0.01 pts (Bonf interval spans zero on fp, targets, carries), reproducible on fresh bytes; the +0.0075 fantasy shift is indistinguishable from an information-free refit perturbation (P - placebo CI95 spans zero); D-209's model-free gain is real and is not visible in the served line. DOES NOT SURVIVE: 'redundant' for carries as a stated fact (corr borderline and concentrated in 2025; the walk-forward test cannot see a signal that only appears in the test season); the shuffled-label control as written; the framing 'MEASURED NOTHING end to end, and the reason is now known' (nil is measured; the reason is measured for the perturbation half only). UNTESTED: the other 33 H-021 windows; T (needs MATCHED_BOTH_BALL); grading on issued aar_forecasts; whether the 2025 carries corr replicates in 2026.

Next action

(1) Treat the record's screen ('model-free gain + corr(x, shipped residual) needs no refit') as valid only with a season-stratified corr and a game-cluster interval, and add the placebo refit (one script, ~40s per fit, 3 seasons in parallel) as the information-free control for every H-021 window instead of a shuffled-label MAE. (2) Score the 2025 carries corr on 2026 weeks once 4+ weeks settle (aar_forecasts vs aar_outcomes) before any 'redundant' claim on carries. (3) Bind a byte-identical snapshot in input_artifacts if this record is ever promoted; the execution snapshot is deleted (record retains rows and hashes).

Source provenance and publication scope

Owned research record: research/scientist/experiments/2026-09-26T2001Z-referee-h021-p-ordered-shadow.json

This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.