Review of 2026-09-23T0611Z-h109-concentration-denominator
Owner lane: Not recorded
Reproduction
REFUSED. Bound bytes verified: record dab83d4b..., protocol 6e0be4ba..., runner 3da30295..., result 6750c1f4..., share.py 37758ac8... (matches HEAD). concentration.py CHANGED: bound 79c79513..., disk/HEAD 7d541eb0... (C-13 d7e40a23 added a timezone('UTC', now()) AS ingested_at column to the team_concentration SQL; values unchanged, verified by replica max abs diff 2.2e-16 on every key in the reproduction). Not bound: the runner freezes `git archive HEAD`, which is now 276ecbcd, and HEAD is 77 playerweek/ files past aa0f6737; the source snapshot sci-h109.duckdb is gone and its hash is in no field; baseline.status NOT_BOUND (refit per fold, PW_FITCACHE=0, no release/model/source-SHA); information_set NOT_VERIFIED. The reproduction pins the tree to aa0f6737 but runs on a new warehouse vintage: shipped MAE 3.72013 vs 3.71988 (+0.00025), so warehouse rows moved slightly even for 2023-2025. That vintage drift alone moved the MAE lower bound from +0.00018 to -0.00036.
Finding and verdict
WEAKENED
Review scope
UPHELD: the premise (squad-denominated WR target 0.228-0.234 vs shipped 0.384-0.394; target-solve fallback 0.254->0.040; median solved target exponent 1.93->0.93) and the direction of the consequence (squad MAE +0.0155 pts, squad-T split +0.0018, oracle_serve not better than shipped) reproduce on a fresh snapshot with the tree pinned to aa0f6737: squad MAE +0.01547 vs record +0.01557, squad-T +0.00182 vs +0.00188, squad_shuffled +0.01948 vs +0.01942, oracle +0.0092 vs +0.0094; shipped_again 0.0, manual vs score_arm 5.0e-5. WEAKENED: (1) the record's REFUTATION CLAUSE (squad worse on fantasy MAE with Bonferroni-3 excluding zero) rests on a lower bound of +0.00018 and does NOT survive on the same code and rows: reproduction Bonf3 [-0.00036,+0.03186]; week-clustered [-0.00030,+0.03281]; team-season-clustered [-0.00217,+0.03397]. The squad-T interval spans zero at both HEAD runs in the record itself, so 'worse on the squad-T split by 0.002-0.004' is a claim about afb1d233 only; at the headline tree it is +0.00188 [-0.00023,+0.0039]. The evidence is 'no gain shown, harm not established', not 'harm established'. Prereg P3 (nil, |d|<0.005) is not confirmed either: point estimate 0.0155 exceeds 0.005 and the interval spans both 0 and 0.005. (2) The MAE cost is not a team-identity effect: squad_shuffled costs the same (+0.0195; squad minus shuffled -0.0040 [-0.0126,+0.0043]). It is carried by 2024 (+0.030) and RB (+0.034); 2023 is -0.0006. The record says 'worse in both code states' but the two runs are the same rows and the same season fold, so it is one sample twice, not a replication. (3) The mechanism sentence 'the target is operating as a fitted sharpening knob, not a measurement of concentration' is not identified by the oracle arm: oracle_serve fits calibration on the shrunk trailing squad target (KEEP=0.35) and serves the unshrunk realised SEASON leader share, so fit and serve targets differ by design, and a season leader's share is not the expected share of the week's model-picked top WR that alpha_for_subset solves for (the H-202 estimand question the record defers). Oracle worse-than-shipped therefore does not separate 'wrong-unit target compensates downstream' from 'season-leader share is the wrong estimand at either denominator'. (4) The decision endpoint (start/sit) is not measured, as the record says. Do-not-ship of the relabel alone stands on no demonstrated gain. Not a PROMOTION.
Source provenance and publication scope
Owned research record: research/scientist/experiments/2026-09-25T1921Z-referee-h109-concentration-denominator.json
This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.