Does the one-banded-residual-distribution-per-quantity calibration (props.conformal_cal: bands by projected level, pooled over positions) over-price RB receiving OVERs by 10-13 points and QB rushing-attempt OVERs by 14 points, and does a per-position (Mondrian) calibration remove it?
Owner lane: s-observation
Question
Does the one-banded-residual-distribution-per-quantity calibration (props.conformal_cal: bands by projected level, pooled over positions) over-price RB receiving OVERs by 10-13 points and QB rushing-attempt OVERs by 14 points, and does a per-position (Mondrian) calibration remove it?
Prediction
Frozen in prereg/2026-09-23T1317Z-h181-prop-position-pooling.md before any run: REFUTED on magnitude for RB receiving -- bias_A in [0.00,+0.08] for both RB receiving cells; QB carries bias_A in [+0.03,+0.12]; per-position bands reduce |bias| in all three cells and are Brier-better or equal. Refuted if any RB receiving cell has bias_A >= +0.10 with 99% CI excluding +0.08, or QB carries bias_A < +0.03.
Finding and verdict
TESTED -- H-181 MAGNITUDES REFUTED; POSITION HETEROGENEITY REAL BUT TWO-SIDED. RB receiving OVERs over-priced +0.045 (receptions) / +0.043 (rec_yards) at a line at our own number, not +0.10..+0.13; QB rushing attempts -0.006 [99% -0.052,+0.037], not +0.14. A per-position (Mondrian) band moves RB receptions by +0.030 (perm p 0.015 against a 200-shuffle calibration null); RB rec_yards' +0.015 is inside calibration noise (p 0.15). The pooled band's larger error is on TE, where it prices OVERs LOW by 0.036 (rec_yards) -- selected post hoc. Per-position is Brier-better on the three primary cells but Brier-worse or equal on WR/TE receiving, so it is not a straight replacement. QB carries cannot be banded per position from one calibration season.
Reasoning
Kept the drawn row. Mapped the claim onto props.conformal_cal (bands by level, pooled over positions) and reused H-180's production-path harness, adding a Mondrian arm.
Next action
(a) The frozen A-P interval method is wrong for calibration contrasts: any future calibration comparison must resample calibration rows (or use a shuffle null), not only scored weeks -- H-180's A-B CI [+0.023,+0.024] has the same defect. (b) A per-position calibration needs >1 season of calibration rows to band QB carries/rush at all; test a 2-season calibration window with position strata. (c) TE under-pricing (-0.036) and WR rush_yards over-pricing (+0.092) deserve their own frozen run, since they were found by looking. (d) Real book lines, not floor(cond)+0.5.
Source provenance and publication scope
Owned research record: research/scientist/experiments/2026-09-23T1330Z-h181-prop-position-pooling.json
This public reading view includes authored question, finding, review, reasoning and next action fields. Raw measurements, commands, logs and local paths are withheld. The record ID clock is not proof of completion or deployment.