27 Sep 2026 · 10:06 AM MTUpdated…
Authored research record

Research output

← All research outputs

refereeReviewed

Review of 2026-09-23T1958Z-h114-qb-regime-judgement

Owner lane: Not recorded

Finding and verdict

WEAKENED

Review scope

SPLIT. UPHELD: every number the runner emits reproduces (64 rows, identical key set, summary identical, only row order and 10 float fields at <=1.8e-15 differ, so not byte-identical because the runner's GROUP BY has no ORDER BY); classes sum to 61; J1-J4 all have offense_pct>=0.5 and J0 none (36 started / 25 not); the 15 J1 men are all incumbent==self; the prediction was scored WRONG honestly. DOES NOT SURVIVE AS WORDED: (a) the registered falsifier verdict '40 of 61 regime QB misses carry no signal of a change' is an ARM-UNION artefact: the 61 entities are the union of four arms' regime labels (168 arm-rows; the record says 171); per arm J0+J1+J5 is S 11/29 (38%, falsifier does NOT fire; J0 = 2), SB 31/46 (67%), SBQ 31/46 (67%, the AAR default arm, fires), SQ 30/47 (64%); J0 is 23 of 25 in 2025 and 2 in 2024; (b) the post-hoc 'a knowable signal the model does not read' (J1, 8.54 vs 19.17) is conditioned on having started, and the unselected same-signal population shows the model already prices it as a mixture (see 3_METRIC/4_SELECTION); (c) 'regime' in attrib.py is the largest of four decomposition components on an out-of-80%-interval miss, not 'change of starter'; the record's 'regime QB errors' is a population of largest-component misses and its J0 class (over-projected non-starters, the biggest) is a stack/calibration term, not a missed start.

Source provenance and publication scope

Owned research record: research/scientist/experiments/2026-09-26T1258Z-referee-h114-qb-regime-judgement.json

This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.