27T1459Z-referee-h207-band-ordering
UPHELD
Recent findings, experiments and independent reviews from the research lanes. Open a record to read its authored conclusion and next action. Activity counts measure output, not scientific quality or live product gains.
“Today” uses Mountain calendar dates. Times come from record IDs in UTC and are shown in Mountain time; they are not verified completion times. A record identifies its lane, not which paired scheduled job alias ran it. Unknown owner lanes stay unassigned. 139 total records = 80 lane findings + 59 independent reviews. A finding can mix positive and negative conclusions, so its authored verdict appears in full rather than a guessed pass/fail label. Adoption and a live shipped gain require separate release proof. 0 unreadable eligible records were omitted.
UPHELD
UPHELD
TESTED. Of kicker fantasy MAE 3.722 (market-anchored walk-forward 2023-25, 1,634 kicker-weeks), actual FG ATTEMPTS at the fitted conversion is worth -1.607 [-1.814, -1.395] (43%); actual conversion and distance at the FITTED attempts only -0.483 [-0.616, -0.365] (13%), so attempts carry 3.3x the conversion arithmetic and the surprise…
SURVIVES: exact, field-for-field reproduction of the analyze step from the frozen preds_wf.csv.gz and a freshly-taken warehouse snapshot; production entry point used correctly (wf_std.py -> PM.fit strict season<S -> share.feature_sql -> PM.lines_for); no leakage in the fit window or the anchor-recentring path; MAE and MSE kept separate…
Every falsifiable number in the record reproduces exactly, and the finding survives a harsher shuffled-outcome attack than the frozen protocol itself required. Record's own verdict stands unweakened: on 816 REG games 2023-25, the market-free HEAD board's summed margin is +1.08 [+0.68,+1.47] pts/game worse than the close (Bonf-6), 53% as…
TESTED. The capture-timing measurement holds as registered (wk2 designated miss 114.5 = 5.8% of 1,987.2, all captured after the build and 108.1 of it after kickoff; wk1 0), but the row's causal counterfactual is refuted: the overlay reads week 1 by a literal, so an on-time capture reaches 1 of 50 week-2 and 2 of 70 week-3 designated men.…
Every falsifiable claim in the record survives reproduction one day after binding. The record's own verdict stands unweakened: REFINED -- a small, real, walk-forward rush-anchor gain (-0.041 team carries/game, Bonf-2 [-0.078,-0.005], 7/7 seasons negative) that does not propagate past this harness's +-0.004 pt/row resolution to the player…
PREMISE CONFIRMED AND MEASURED. The 0.45/1.75 pair is live on every K/DEF row of the warehouse forecast table and reached the served K row in week 2. Held out 2021-2025, it breaks its floor 16.6% (K) and 15.8% (DEF) of team-weeks against a nominal 10%, and its ceiling 12.1% / 10.7%. A walk-forward residual pair (C) beats it on pinball: K…
SURVIVES: exact byte-for-byte reproduction of the analyze step from the frozen learned_baseline and all four bound input artifacts; the production entry point (wf_std.py -> PM.fit -> share.feature_sql -> PM.lines_for) is used correctly with a strict season<S fit window and no ALL_FEATURES/SELECT * defect; the placebo control is a…
As scoped by the record itself: 0216Z weak joint (1) CLOSED as a sensitivity result (denominator moves the served B1-B0 conclusion by at most ~0.05 pp and never flips its sign or opens its 95% interval to 0 across 9 arms); a NEW finding that the served order carries a level jitter of about +-0.05 pp, which is the same size as the…
SPLIT. Coverage half CONFIRMED (9 of 32 kickers priced on the 2026-09-20 board, all rostered; DEF prices 32). 'No ordering skill' REFUTED as literal: the served kicker line orders kickers at within-week Spearman +0.144 [+0.109, +0.178] on 2019-25 walk-forward, exceeding all 200 shuffle draws, and every bit of that order is the team…
As scoped by the record itself (verdict REFINED, prediction refuted by its own clause, f-specific flex_sym interval spanning 0). The design-specific falsifier -- a 200-draw permutation null of f pushed through the identical served pipeline -- does not remove the effect: B1-B0r +0.299 pp against a null mean -0.005 pp (sd 0.030, max…
The measurement reproduces byte-for-byte and its own falsifiers pass; what does not survive is the attribution. f is one weak proxy of a season-level under-read, not the under-read: season-to-date fantasy production (g) as the added term gives +1.16 pp on close and +1.50 pp on flex_sym, and once g is in the model f adds nothing (close…
TESTED. The premise as worded is false since 2026-09-23 and the fill persists through tonight's refresh (0 mismatches vs pbp over 5,344 team-weeks, 2016-2026). The 'wrong sign' reproduces but was never identified: bias -0.55 [-2.02, +0.84] with stops vs +0.51 [-0.90, +1.83] without. What the zeroed stops actually did was flatter DEF MAE…
UPHELD as scoped by the record itself: a real sign-stable slope step (2023-25) and a P calibration slope above 1 on rush; the MSE value of fixing it is UNRESOLVED (interval spans zero, lower edge -1.2% of MSE), not shown to be zero. The record's own caveats (post-hoc calibration interval with Bonf-8 lower edge 1.00, team-game grain…
WEAKENED
TESTED. The COALESCE end is a real fit/serve encoding mismatch -- playprob/qbshare are fitted with NULL=3 and served 0, which they read as 'position average, snaps ignored' (P(play) 0.885 vs 0.716 on 6% of 2023-25 roster-weeks). But correcting it in either direction measured WORSE on mean-sensitive accuracy: the frozen remap wins MAE on…
WEAKENED
TESTED. H-110 as registered is REFUTED: on the served row the WR room is as concentrated as a draw from its own p, 2025 +0.0019 [-0.0066, +0.0101], pooled 2023-25 -0.0009 [-0.0056, +0.0038]. But that endpoint cannot test the premise in H-110's title: HHI is invariant to WHICH man holds each share, and scoring it on projections shuffled…
REFUTED AS REGISTERED, BY 0.05pp. Shipped-band ceiling exceedance for men outside the room (8.1% of played rows) is +3.13pp over full-room men at matched level, Bonf-12 [-0.05, +6.08]; same sign in 2024 (+2.1) and 2025 (+3.1, held out, Bonf-12 excludes zero) and on arm S (+2.9). The version the data carries: it is spread, not bias…
WEAKENED
H-151 REFINED. As worded it is REFUTED: for a dressed man production serves P(played | dressed, cell), not P(dressed), and on the bench it is served too LOW (-0.09 to -0.13), not too high. The claim the data carries: the table's population changes meaning at 2019 (no INA before), so it is fitted on a mixture of P(dressed x played) and…
WEAKENED, not refuted. UPHELD: the zero of a points gate is not the zero of the bet: at the taken prices the breakeven mean CLV is about +0.39 pts under an efficient-close assumption, and a t95 lower bound on CLV points at n < ~88 can be cleared by a population that loses money; on the 35 stored rows every conversion has a negative point…
H-149 AS STATED IS REFUTED on the frozen bar: outside the P(play) mixture at the bottom two bands, 0 of 72 band x {position, opponent, game total} x endpoint cells shows two opposite-sign cells at Bonferroni-96, discovery or held-out. 'Every per-tier readout' is false. Prediction: S2 partly right (bands 0-1, not every band: dnp rows…
REFINED, CONFIRMED AND EXTENDED. (1) H8 binds the 2026 board's own window: fit season<=2025 moves 12,225 and 13,721 of 13,721 rates under two reorders (max 2.72), not only the <=2022 invariant window. (2) It is exactly the multi-position ids: relabel those 32 players (1.1% of played-row targets and carries) and both reorders move 0 rates…
WEAKENED, not refuted. UPHELD: the primary target-share recency residual on the targets endpoint is real out of selection (2019-22 and 2023-25, player-clustered t 4.2-4.5), and small (bound 0.0011 targets/row; 0.0053 at weeks 2-4); target-share and snap-share hl2/hl4 -> targets survive Bonf-7. NOT UPHELD: 7/7 cells (4 of 7), carry-share…
H-259's premise REFINED AND ITS PREDICTION REFUTED AS CAPTURED. The structured feed was NOT no-later-than-text: S first or tied on 3% (wk1) and 2% (wk2) of elevations either named, against the >= 80% predicted; the cause is capture schedule, not source latency (P1). Text named 79-81% of labelled elevations pre-kickoff, but the…
REFINED: sol's stop was correct as a gate and too wide as a scope. H8 binds ZERO market-grain quantities: the published game margin/total, the wagers decide-week writes and the CLV instrument on chain-report never load efficiency.fit, and a physical reorder of the one table they read moves 0 of 2,911 outputs while a +7 score perturbation…
WEAKENED, not refuted. Upheld: nil for td_oe_l8, hurry_delta_l8 (committed definition, and current definition as sensitivity) and own_scramble_rate_l4 on the E2 correlation, and E1's direction. Not upheld: (a) 'MEASURED NOTHING' for the three season shares - target_share_season and snap_pct_season (hl2) on the targets endpoint carry a…
REFINED STOP (the H8 stop stands for exact replication; its scope is narrower than stated). H8 at consumer grain, 2023-25 held out: reordering identical rows moves 6,134-6,155 of ~6,840 fantasy forecasts, 5,376-5,423 target and 4,648-4,746 carry lines per season, but by mean |dFP| 0.028-0.036 pts (p99 0.19-0.29, max 1.97), mean…
The premise holds and is stronger than filed once game script is removed (PROE r +0.37 same-HC vs -0.01 new-HC over 19 seasons). Shrinking the season-crossing throw-rate input toward the league for new-HC teams cut team throw-attempt error in weeks 1-3 by 0.24 attempts [99.375% 0.07, 0.43] on held-out 2019-2025, in every one of the 7…
UPHELD for 'the 16 EWMA-hl2 recompositions of player_pbp_l4 add nothing measurable to the shipped fantasy line' (E2 linear, winsorised linear E3, and a shallow nonlinear check), scoped to rows where both pools exist. NOT UPHELD as mechanism language: the E1 gain over flat4 is pool depth, not recency. Not a claim about the 17 unscreened…
STOPPED: H8 VIOLATED; H-194 remains untested. This repeats the already filed H-274 physical-order defect.
SURVIVES the falsifier: exact replica, production-path P vs shipped nil at about +-0.01 pts (Bonf interval spans zero on fp, targets, carries), reproducible on fresh bytes; the +0.0075 fantasy shift is indistinguishable from an information-free refit perturbation (P - placebo CI95 spans zero); D-209's model-free gain is real and is not…
STOPPED: H8 VIOLATED; no market hypothesis tested
SURVIVES: direction and the ~1.2-1.3pp product accuracy cost / +0.15 pts regret on pair sets not conditioned on either arm alone, under both the record's previous-played-row arm and the production involved-row selection; post-fit arithmetic reproduces byte-identically. DOES NOT SURVIVE: the frozen primary's exclusion of zero as a stable…
STOPPED: invariant H8 VIOLATED; H-194 untested. H-274 and BLOCKED.md already name and localize this defect.
WEAKENED
SURVIVES: values reproduce exactly (0 of every levels/contrasts entry differs; n 20,066, 1,632 clusters, meta identical) and the direction of the effect. Ceiling exceedance is higher on a stale served row than on the target row, in every season, at RB/WR/TE, with no QB effect. DOES NOT SURVIVE: (a) the headline number as 'the board's'…
SURVIVES: the reproduced values (every family endpoint within drift; no sign or exclusion flips), the null on weeks-1-3 width (E2 spans 0 on all five), rush_yards centre E>L in 5/5 seasons, and the harness controls (K1 in-sample <=0.0072, K3 unchanged). DOES NOT SURVIVE: (a) 'H-185 REFUTED AS FILED' for the QB throwing-yards centre and…
SPLIT. SURVIVES: the direction of the headline. Blending the centre does not hurt the SHIPPED_CURVE band and refitting the curve on the blend does not help it, in the one genuinely held-out season (2025): secondary B-refit - B-ship pinball +0.0013 [-0.0053, +0.0076], cover -0.0028 [-0.0061, +0.0004]. The shuffle falsifier fires (+2.29…
SPLIT. UPHELD: the frozen B endpoint's direction. Scaling projected non-QB1 QBs' rush fields by P_exit before _reconcile makes fantasy points worse in both seasons, and worse on MSE as well as MAE, so it is not a mean-vs-median artefact; the shadow adds +0.5-0.6 RB carries a team-game of bias to a chain whose RB carry bias was ~0 (+0.05…
SPLIT. UPHELD: every number the runner emits reproduces (64 rows, identical key set, summary identical, only row order and 10 float fields at <=1.8e-15 differ, so not byte-identical because the runner's GROUP BY has no ORDER BY); classes sum to 61; J1-J4 all have offense_pct>=0.5 and J0 none (36 started / 25 not); the 15 J1 men are all…
SPLIT. UPHELD: every number the runner emits reproduces byte-identically (wk2 21 men / 36.0 pts, wk1 25 / 28.5 under the runner's join; A/B/C/U cells; E3 means; the six post-hoc share>=0.05 men and 29/35 below it) and the record's own verdict text ('mechanism NOT established', 'PARTIAL') is honestly scoped. DOES NOT SURVIVE AS WORDED:…
SPLIT. UPHELD: the primary claim (E1 carried 0.38, E2 pooled residual +0.082/pt, CI excludes 0 at every position and every season) reproduces byte-identically and survives the level-control, the p_play>=0.9 selection restriction and a correctly split two-sided Bonferroni-26. DOES NOT SURVIVE AS WORDED: the E3 'mix signal' sentence (fixed…
TESTED -- H-158's cost claim REFUTED by its registered rule. The premise holds: on the live board's row, share describes the final three games and volume/snap the three before (fixture: shipped volume matches the Wednesday window on 17.6% of rows, snap on 5.6%). Aligning them buys nothing on the decision: SVN - S -0.10pp [Bonf-4 -0.40,…
Two of the record's three claims stand and one does not. STANDS: (a) coverage, (b) the mechanism, refined. DOES NOT SURVIVE AS WORDED: 'decision cost nil in weeks 1-2', because its only test is a preregistered threshold (24th-best flex among rostered played men) scaled to a league this is not. Sean's league is 10 teams with 70 non-QB…
MEASURED NOTHING, and the premise is void. H-189 locates a win that does not exist: on the stored held-out ledger SB beats S by +0.007 pts [-0.008,+0.024], so its 111% is a ratio with a nil denominator (F1 share 0.35, 95% [-2.30,+3.57]; a random floor gives [-0.32,+1.48]); the figure reproduces only when the floor is ranked by SB's own…
The record's numbers reproduce byte-for-byte. What does not survive is the headline reading of the RB row: (1) its interval excludes zero only under a game-cluster bootstrap, while the treatment (a defence's trailing-points tercile) is nearly constant within a defence-season, so the independent unit is ~96 defence-seasons, not 816 games;…
H-292 CONFIRMED HELD OUT. On 2019-22 (seasons H-207 never scored), across 41,558 close same-position RB/WR/TE pairs (lines within 1.0, both >= 6, played=1 SELECTED), the man with the higher sim_world centre outscores the other 54.0% of the time, against 51.5% for the man with the higher published line: sim minus line +2.53 pp [99% +0.99,…
H-207 FAILS ITS BAR, and the band carries no SPREAD-ordering skill: the width is a location signal. E1 (top tail) +3.44 pp [99% +1.86, +5.04], outside the shuffle null; E4 (underdog rule, synthetic head-to-head) -0.085 pp [-0.174, +0.008], inside the shuffle range and nil. Among close RB/WR/TE pairs 2023-25 (30,069, played=1 SELECTED)…
MEASURED NOTHING. Geometric recomposition of the ten team windows (own + opponent copies, both sides of every game, every consumer of team_week at once) moves fantasy MAE +0.0048 [Bonf-3 -0.0088, +0.0187], targets -0.0030, carries +0.0023, all spanning zero, on 20,523 player-weeks 2023-25 with the replica exact. The whole block is…
Headline survives, three secondary claims do not. SURVIVES (non-exact reproduction at HEAD d8c149dc on a fresh snapshot): zero - shipped pooled +0.0137 [99% +0.0023, +0.0257] (record +0.0156 [+0.0038, +0.0278]), interval excludes zero, point above the frozen +0.005 refutation line; RB +0.0303 [+0.0046, +0.0578] is again the only position…
PARTIAL (frozen rule: model b_R Bonf-6 upper < 1, A2-A1 spans zero). The line's week-to-week revisions are 15% too large (b_R 0.853 [Bonf-6 0.824, 0.887], 47,223 player-weeks 2019-25), not the filed 22%, and the largest 5% are not worse (0.846, filed 0.65). FantasyPros underreacts (1.051 [1.003, 1.101]). On the board's metric a revision…
H-289 CONFIRMED ON THE MOVED PREMISE, and the prereg's arm B MISSES its bar. At HEAD (no market anywhere in the line) the summed board is +1.08 [+0.68, +1.47] pts/game worse than the close against the result on all 816 REG games 2023-25 (2023 +0.87, 2024 +1.10, 2025 +1.26 -- fable's +1.03/+1.32 reproduce on the unselected rows), with 53%…
Narrow. Upheld: every number in the record reproduces to the printed digit on a fresh snapshot at HEAD (b-a +0.0037 [-0.0057,+0.0131], c-a +0.0013, d-a +0.0077, X-a +0.0067, team pass/rush contrasts, per-position, per-phase, per-season); arm b is leakage-free (poisoning every team_week row at or after week w changes served() on 0 of 1728…
REFUTED on its model half; football half reversed on rates. Held-out 2019-2025 walk-forward: the production line does NOT under-project movers' targets (+0.013 [-0.027, +0.056], inside placebo), not by week, season, rostered population, rate, share or room. The served 2026 board shows -0.18 / -0.10 on 56-68 movers, intervals spanning…
REFINED. Anchor: rush recomposition (nested ewma4-8) beats shipped l3 by -0.041 team carries/game Bonf-2 [-0.078,-0.005], 7/7 seasons negative, 0.7% of rush MAE; pass nil (-0.012 [-0.042,+0.019]). Player line: MEASURED NOTHING (fantasy -0.0002 [-0.004,+0.004] Bonf-6, carries +0.0006, targets -0.0004), resolution +-0.004 pts. H-021 end to…
Narrow. Upheld: (1) rapm_theta and uplift_cate are 1.000 vs 1.000 played vs not-played inside each of RB/WR/TE and have zero QB rows, so the pooled 0.90 vs 0.46 is position mix; (2) regime_gate and kalman_share presence equals the SQL rule on 23935/23935 keys AND equals the output of the real builders re-run on the snapshot; (3)…
Narrow claim only: at HEAD f6511555 latest_rows(INVOLVED_WEEK=True) selects a week with a player_pbp_l4 row whenever the player has one at or before the selected week (0 stranded on the season-2025 board, 633 rows, and on today's season-2026 board, 450 rows); INVOLVED_WEEK=False strands 115 (2025) / 33 (2026). NOT upheld as stated in the…
REFINED (premise). 0216Z weak joint (1) CLOSED: the level denominator moves the served conclusion by at most +-0.05 pp and never changes a sign or a 95% verdict on B1-B0. Added: the served order carries a level jitter of about +-0.05 pp, so the B0r-B0 recalibration artifact is sign-stable but its size is uncertain by 2x. H-021 end to end…
UPHELD and REPRODUCED: the frozen prediction (b-a within +/-0.01, REFUTED only if b-a Bonferroni-2 interval clear of +0.01) survives the attempt; every field of the record's result reproduces exactly (b-a -0.0034 [-0.0104,+0.0037], c-a +0.0000 [-0.0061,+0.0061], X-a -0.00003 [-0.0076,+0.0072]; replica max|diff| 0.0 all seasons), and the…
UPHELD: the numbers reproduce exactly, and 'H-181 magnitudes refuted' survives (RB receiving +0.045/+0.043, QB carries -0.006, none near +0.10..+0.14 on the scored population). WEAKENED: (a) the scored population is all p_play>=0.5 men, while my stratification shows the over-pricing concentrated in the p<0.95 men (RB rec_yards +0.081…
Two halves with opposite outcomes. (A) The WALK-FORWARD half REPRODUCES (non-exact) and is UPHELD: TE targets share ON -9.34% [95 -10.90,-7.83] (record -9.32), OFF +1.06% [-0.66,+2.69] (record +1.08); WR targets ON +4.33% (record +4.33); TE points MAE OFF-ON +0.0810 [99.6 +0.0626,+0.1011] (record +0.0813); all-rows -0.0009 (record…
REFINED (serving). The 0118Z order gain reaches the served product, but a quarter of the served B1-B0 is the blend re-weighting itself, not the share term. f-specific served gain: +0.29 pp flex_sym (edge of zero, as raw), +0.21 pp close 2023-25 (excludes 0). New: the served order is recalibration-sensitive because 10.2% of served lines…
TESTED. H-161's premise STRENGTHENED: the rush-TD slope step is not one season, it is 2023, 2024 and 2025 each ~2 SE above 2016-22, and held out the pooled production anchor is conditionally MISCALIBRATED on rush TDs: calibration slope 1.31 [95% 1.08, 1.55] on 2023-25 (Bonf-8 lower edge 1.00), high-minus-low implied-tercile bias +0.132…
REFINED. H-021 end to end is still MEASURED NOTHING on MAE; on the start/sit ORDER the season-share under-read is real on close calls (+0.50 pp [+0.32,+0.69], regret -0.059 [-0.081,-0.038] pts per call, 2018-25, 8/8 seasons, held-out 2018-22 alone +0.50 pp) and at the edge of zero on the primary flex_sym 2023-25 cell (+0.46 pp, Bonf-2…
The measurement REPRODUCES (non-exact, fresh snapshot) and its falsifiers hold; what weakens is the benchmark identity and the framing. (1) REPRODUCE: the runner text, changed only in snapshot name and OUT path, on a fresh snapshot gives p<0.95 L0 pooled n 4,580 (record 4,577), bias_A +0.0717 [0.0524,0.0880] (record +0.0720…
REFINED, and 23:17Z's headline REFUTED out of selection. P0 on the unscored seasons 2017-18 is +0.011 [-0.051,+0.076], inside Bonf-8 (refutation clause met). The season-share lead is not under-shrinkage: it is the mean line's over-dispersion (slope 0.83 at every week, not an early-season defect) netted against an opposite and larger…
UPHELD as counts and mechanism, on a replay -- NOT as an exact reproduction. M1: props.settle keys actuals by player only (source: props.py:1032-1037, `act.setdefault(pid,{}).update(...)` ORDER BY season, week, so the LATEST played week wins). Independent replay of the 09-23 state (player_week truncated to week<=2, current snapshot, my…
H-138 AS FILED: NOT REACHABLE at the release (0 market rows on every served population); latent under PW_MARKET_ON=1. REFINED: the first-row read also governs the served volume anchor but only as a ceiling/gate (<=2.90 team pass attempts, 4 rooms, order-dependent); the real preseason volume defect is that 163 moved men are priced on the…
REFINED: the 21:17Z week 2-4 lead is under-shrinkage of the season-to-date target share, NOT prior-season carry-in and NOT recency weighting. Frozen primary held (C1 outside Bonf-6 in both populations), but a post-hoc amendment registered before it ran shows C1 is a proxy: last season adds +0.028 given f (inside band), while -f alone…
UPHELD, as a re-execution on a fresh snapshot with the tree pinned to bdd04f82: the FROZEN composite SWAP (order by c = served/p_play, then swap if the pick is INA) is worse than the shipped DISCOUNT on eligible pairs. SWAP-D regret +0.17032 [+0.01965,+0.30662] Bonf-2 (record +0.16203 [+0.01419,+0.29869]), accuracy -0.00782…
H-272 CONFIRMED AND SHARPENED. On today's 35 rows fable's literal falsifier holds (both gates closed, both negative) -- but the EV interval -4.95% [-9.41, -0.49] per unit EXCLUDES ZERO BELOW under every conversion A0-A4, a statement the points gate [-0.75, +0.52] cannot make: the process as run is below breakeven at the fair-close…
H-137 as filed: NOT OBSERVED -- the 'three teams' are priced now, the latest_rows stand-in is withheld by G-TDANC (0/450 rows), and K/DEF keys all 32 teams to our own score by design (544/544). The refined claim the evidence could carry -- that the K/DEF slot fitted on the book but served ours is a fit/serve mismatch worth closing by…
UPHELD (reproduces on a fresh warehouse snapshot with the tree pinned to 914106d5): the mechanism-stage premise (median target exponent 1.7013, WR room +0.0807 of squad target mass, RB -0.0569, TE -0.0226; carry exponent 1.2005, RB +0.0241, QB -0.0198 -- bit-identical to the record's table, it is deterministic), the null-to-small rp_same…
H-130 REFINED AND SUPPORTED on the production builder, stated more broadly than it was filed: forecast.build discards team QB-room volume for any team whose starter carries a curve designation, market or no market. Out: 100% of the room in the designated week (7 of 8 teams), 68% at k=1, 28% at k=2, 8.6% every week after. Questionable:…
CONFIRMED OUT OF SELECTION and LOCATED; still small. Primary +0.041 on 2019-22 (4/4 seasons), all 7 cells outside Bonf-7 in both populations. Concentrated in weeks 2-4 (rho ~0.09 pooled, linear bound 0.0053 targets/row = 0.38% of MAE 1.41) vs week 5+ (0.0007). A 10-bin walk-forward correction is nil (+0.0074 vs shuffled +0.0066).…
UPHELD: Part A enumeration reproduces to four decimals (WR squad gap +0.0483, RB carry +0.0821, group WR +0.0843, RB +0.1051; season leader's own weekly mean 0.2505 WR / 0.5376 RB, so 41% / 58% of the gap is his absences, A3 FAILED as recorded); shipped_again 0.0 max abs; manual vs score_arm 5.0e-5; replica of CC.targets 2.2e-16; fp_l3…
UPHELD: the premise (squad-denominated WR target 0.228-0.234 vs shipped 0.384-0.394; target-solve fallback 0.254->0.040; median solved target exponent 1.93->0.93) and the direction of the consequence (squad MAE +0.0155 pts, squad-T split +0.0018, oracle_serve not better than shipped) reproduce on a fresh snapshot with the tree pinned to…
UPHELD: the served FACTORS band input is not the fitted unit (walk-forward 5.0-8.6% of bands overall, R1 11.5-24.0%, reproduced to within 1-4 rows of 298/482 and 285/465), and feeding it moves pooled fantasy MAE by +0.0006 [-0.0040,+0.0053] (record +0.0008 [-0.0038,+0.0054]), same sign, same interval; the replica and B controls reproduce…
UPHELD: roster3 (six l3 columns over his last three ACT rows) does not beat the shipped projection, and H-224 as filed (age 4+ better by 0.10-0.30) is refuted; reproduced on a fresh snapshot at a later HEAD. WEAKENED on three secondary claims inside the record: 'pooled nil HELD', the age-4+ 'worse than shipped, p=0.065' half that drives…
UPHELD
The enumeration (which tables, which writers, coverage, exposure fractions) survives an attempt to refute it at the current tree. Two secondary claims do not: the served-0-is-out-of-support characterisation (0.0 is an existing ~1% sentinel) and the ~09-29 timing (the selector can flip on the first played week-3 row). Source record…
UPHELD for the question asked: the -3.50 measures the metric (perfect forecaster reads -3.6 to -3.8 in weeks 1 and 2, exact sign-flip p<=1e-3). WEAKENED for two secondary statements: the recorded week-1 magnitudes (B-A +3.64, D +0.13, SEL 2.15) do not reproduce from today's warehouse (exactly 20.09 targets of played-flag drift), and…
MEASURED NOTHING end to end on H-021's last 6 player windows (24 of 35 now nil at this resolution); ONE STABLE RESIDUAL LEAD: season-share recency is under-read on the targets/carries endpoints (rho 0.03-0.05, 3/3 seasons for target share), worth <= 0.0013 targets/row linearly, below any refit's +-0.01 resolution
MEASURED NOTHING end to end on 16 of H-021's windows; the model-free EWMA gain is real (16/16) and does not reach the line
WEAKENED on the rho inference and the trend mechanism; UPHELD only for the perfect-forecaster control (byte-exact reproduction, actuals-only). Source record stays UNADMITTED (legacy, no evidence contract); nothing here is a PROMOTION and no CHANGES-PROPOSED entry is implied.
WEAKENED
UPHELD only for the narrow property named in the bar (both functions' season literal follows LAST_COMPLETE_SEASON). The source record itself stays UNADMITTED (legacy, no evidence contract) and its `no_regression` clause is superseded by CP-M5. Nothing here is a PROMOTION; no CHANGES-PROPOSED entry is implied.
MEASURED NOTHING on the production path, exactly, and that is the result: D-77 GRADIENT_SHARE has zero reach at the release config (0 site calls, 0 of 6,867 rows moved by any of 12 arms), because no production row carries mkt_team_total since G-TDANC (2026-09-15). The positive control that the 09-25 referee lacked works (keyed market,…
MEASURED NOTHING end to end, and the reason is now known. Exact replica, production path, 20,523 player-weeks 2023-25: EWMA(hl=2) on targets_l3 + carries_l3 moves fantasy MAE +0.0075 [Bonf-6 -0.0065, +0.0215], targets +0.0025, carries -0.0003, every interval spanning zero. The D-209 recency signal exists before the model (-0.052 targets,…
WEAKENED
H-130 as worded is NOT OBSERVED on the published object: 0 of 308 designated-QB team-weeks across 3 generations lose any anchor mass beyond 4-dp rounding. But MEASURED NOTHING about the claimed mechanism: no designated QB reached the rebuild, so the test never exercised it. My preregistered prediction (Out teams 0.30-0.70 at k=1) FAILED…
WEAKENED
WEAKENED
WEAKENED
CAUSE IDENTIFIED for the H-021 replica failure: physical row order of player_week reaching efficiency.fit through ANY_VALUE(position) at efficiency.py:182. Not fit non-repeatability, not TEMP shadowing, not the window recompute. Both preregistered predictions held. MEASURED NOTHING about EWMA accuracy. H-021 stays OPEN, blocked on WORK…
WEAKENED
TESTED. The premise holds and its size was over-called. On H-089's own identical pairs, the board's served number orders two startable men worse than the trained number by -0.82pp [Bonf-4 -1.63, -0.01] and +0.12 [+0.03, +0.22] pts regret per pair; H-017's raw arm -0.51pp, interval spans 0. H-089 predicted a further 2-5pp: FAILED on…
MEASURED NOTHING about EWMA accuracy; replica control FAILED. H-021 OPEN, blocked on model-level identity despite H6 green.
REFINED. The conclusion survives and gets stronger; the mechanism is reversed. The line no longer copies the market: G-TDANC switched off D-76 on every production path (R2 0.146 against implied total, 0.989 with the anchor re-fired in memory). What is left is a football-only throwing-TD line that adds no measurable information once the…
WEAKENED
UPHELD
TESTED. The board's floor/ceiling is fitted on a row it does not serve, and it costs the CEILING, not the floor: on the HEAD served row (after CP-H49), SHIPPED_CURVE's ceiling is exceeded 14.3% of the time against 8.3% on the target row every band gate grades -- +5.92pp [Bonf-2 +5.51, +6.34], pinball90 +0.220 [+0.204, +0.237], 20,066…
WEAKENED
TESTED -- H-185 REFUTED AS FILED, REFINED. Weeks 1-3 are not where the prop band is widest (coverage early-minus-late -0.017 to +0.040, all spanning 0; QB residuals narrower early) and not where the QB throwing centre runs highest (-9.6 yd, lower early in 4/5 folds). One piece survives at 95% but not at the frozen family level: the…
TESTED. The centre mismatch H-184 names is real in code and costs the band nothing: serving SHIPPED_CURVE around the blended centre IMPROVES quantile loss by [-0.0417, -0.0516, -0.0315] pts (2025, held out) and [-0.0487, -0.0652, -0.0345] (2024), and refitting the curve on the blended centre adds nothing ([0.0013, -0.0053, 0.0076]). What…
UPHELD
TESTED. H-145's premise refined: (a) 'empty for the live season' is stale for completed weeks (nflverse backfills them) and true only for the target week, where serve is the sole source by construction; (b) 'different vocabulary' is real (79% of labelled rows disagree at K-90m) but costless on 2026 wk1-2 -- all 141 serve-only Out rows…
TESTED. alert=False on every lost firing with rows is CONFIRMED and current. The 'T-90 inactives in hand' premise is REFUTED at the source: ESPN summary returns at most 5 designations per team. Decision reach: 8 of 30 games (09-13 17:00Z) had the warehouse blind for 84.8 of 90 window minutes (lower bound: the drain has no fold clock, so…
TESTED. Premise as worded is stale (8,057 rows exist). The substantive defect stands in a different form: for 2026 weeks 1-2 the lock ledger is a post-hoc backfill by a later commit, differs from the stored board by a one-signed -0.28 to -0.61 pts mean, and carries no sigma at all. Any CLV/coverage grade read off `bets` for these weeks…
REFUTED as registered (bar 1.5 missed: 2.19), and the registered falsifier FIRED: the 10 clean teams sit at +2.885 before and after, so the pool level is part of the board's RB gap. The board shadow does lower the RB gap to prior-season actual (-0.75 [-1.38, -0.20]), but C1 (same mass off the QB1) gets two thirds of that (-0.50), so the…
MEASURED NOTHING FOR THE FORECAST LAYER, AND H-114'S RESCUE CLAIM IS VACUOUS. 22/22 raw<=MIN_RAW rows are share-stage zeros (touches <= 0.0082); volume and efficiency zero none, and the efficiency negative fixture (rates forced to 0) classes 469/469 as EFF, so the classifier could have said so. The stack the guard discards is 0.54 pts in…
PREDICTION WRONG ON THE MECHANISM, MEASURED. 215 of the 222 all-zero lines on the HEAD board replay (97%) are D-95's off-roster stamp (confirmed_active=0: no player_week 2026 on_roster=1 row, matched by resolved name), not an empty row and not p_play. The name key and the player_id key agree on all 215, so there is no crosswalk defect…
PARTIAL / EVIDENCE FOUND, CAPTURE_ONLY_FACTOR_NOT_ADMITTED. Captured pre-kickoff language named 42-44 of the 46 graded zero-line men (98% of their 64.5 pts), and an elevation cue near the name raises P(play) among zero-line rows from 0.09 to 0.45 (Fisher p 1.6e-9, post hoc). The largest single loss (Wentz 19.2) was language-consistent:…
FALSIFIER FIRED as registered: 40 of 61 regime QB misses carry no pre-kickoff judgement signal of a starter CHANGE (25 are non-starters over-projected, 15 are a continuing backup starter). J4 (silent absence) was predicted largest and is the smallest started class. Disclosed post hoc: 15 of 36 started misses are a backup who had already…
PARTIAL. Headline count REPRODUCED (21 on wk2; 25 on wk1). Cost measured: 1.25% / 1.67% of graded half-PPR points, 40% of it (Wentz + White) in two men. The mechanism clause ('a zero share') is NOT established: the registered falsifier fired on inputs, but the class map could not tell a zero model share from a later-stage zero.
TESTED. The '6% of football's slope' premise does NOT reproduce at HEAD: the board carries 38% pooled (42% within-season) -- QB 69%, WR 38%, TE 27%, and RB carries the WRONG SIGN (board -0.024 against football +0.058). The residual slope is +0.082 [+0.067,+0.097] fantasy points per point of game total per player-week, CI excludes 0 at…
TESTED -- PREMISE CONFIRMED, DECISION COST NIL IN WEEKS 1-2. The rostership gate is the whole mechanism: every not-rostered rookie who played had no row (0/84), and rookie coverage was 17% and 11% against 96-97% for everyone else. But the gap has not cost a decision yet: no uncovered rookie scored a flex-starting week (best 14.4 half-PPR…
MIXED, PREDICTION WRONG ON LOCATION. The level-only ceiling does not under-cover receivers against soft defences (WR +0.1pp, TE +0.3pp, both inside +/-2pp). It does on running backs: +2.7pp [+0.1, +5.2] more ceiling exceedances against the top tercile of RB-points-allowed than the bottom, at the same projection level. Keying the curve on…
TESTED. The post-flip coach pair costs +0.016 pts/player-week pooled [99% +0.004, +0.028], concentrated on RB (+0.029) with QB unresolved (+0.032, spans zero); it lowers the within-week rank correlation from 0.7206 to 0.7188. Carry-forward (lag) recovers all but +0.003 (interval spans zero), so the cheapest writer is a coach-keyed…
MEASURED NOTHING at the player (prediction held): served season mean +0.0037 [-0.0057,+0.0131] pts/wk; hindsight season mean +0.0013; league mean +0.0077; shuffled team +0.0067. At team volume the same block is worth 0.20 throw attempts over a random team's, and the served swap is still neutral (+0.034 throw attempts).
TESTED. rapm_theta/uplift_cate: position mix, exactly (1.000 vs 1.000 within RB/WR/TE). regime_gate and kalman_share presence: 100% explained by each builder's rule. Frozen control FAILED on RB kalman (0.416 < 0.5), so the confirmatory reading is void by the row's own terms; counts filed as discovery. Defence: joins clean, spine is where…
CLOSED -- premise dead at HEAD. The defect H-200 names was fixed by D-399 in latest_rows and the invariant holds on the current served board (0/633), with a positive fixture showing it would catch 115 if the flag regressed.
MEASURED NOTHING (NIL, prediction held). Serving the league-mean opponent costs -0.0034 [-0.0104, +0.0037] pts/wk; the real opponent's block is worth no more than a random opponent's (-0.00003 [-0.0076, +0.0072]) or a hindsight season mean (+0.0000).
TESTED -- H-181 MAGNITUDES REFUTED; POSITION HETEROGENEITY REAL BUT TWO-SIDED. RB receiving OVERs over-priced +0.045 (receptions) / +0.043 (rec_yards) at a line at our own number, not +0.10..+0.13; QB rushing attempts -0.006 [99% -0.052,+0.037], not +0.14. A per-position (Mondrian) band moves RB receptions by +0.030 (perm p 0.015 against…
TESTED -- PATH-DEPENDENT. H-174's claim that the error POOL_TGT corrects 'is not there' is TRUE on the in-season path (TE targets share OFF +1.1% [-0.6,+2.7], ON -9.3% [-10.9,-7.8]) and FALSE on the board path at HEAD (OFF +10.9% [+4.1,+18.0], ON -0.6%). By point MAE ON wins TE on both paths (WF +0.081, board +0.140 OFF-ON) and loses WR…
TESTED -- MECHANISM REAL, MAGNITUDE AND SHAPE REFUTED. The shipped prop path over-prices the OVER at p<0.95 by +0.072 [+0.053,+0.089] at a line at our own number -- below H-180's 9-30 floor pooled, but carries +0.120 [99% +0.072,+0.167] and rush_yards +0.079 reach it. Only +0.024 of that is the band/division mismatch H-180 names, and the…
TESTED -- premises 1 and 2 CONFIRMED, premise 3 CONFIRMED AT HALF THE FILED SIZE. (1) Called on today's board, props.settle grades 1,111 week-1 lines against week-2 stats: 88.1% get the wrong actual and 47.3% (526) get the wrong over/under, so any hit rate, reliability curve or the item-1.78 bet gate (tests/test_prop_bet.py step 4)…
TESTED -- REFUTED, WRONG SIGN. Ordering by the conditional line plus a post-list swap costs +0.162 regret per eligible decision (Bonf-2 [+0.014,+0.299]) and does not help accuracy (-0.7pp, spans 0). On this room and this model the served p_play x line is the better start/sit number. That does NOT show the discount is right in principle:…
TESTED -- REFUTED AS A FIX, PREMISE CONFIRMED AT THE STAGE, INERT DOWNSTREAM. Room-preserving sharpening with the shipped exponent is nil on fantasy MAE (-0.002) and SQ_T (-0.0008), both Bonf-6 spanning 0. Its one clearing gain (SQ_C -0.0037) is reproduced by shuffled room labels, so it is not attributable to rooms. Re-solving the…
TESTED. Premise CONFIRMED (weekly leader share exceeds the season figure by +0.048 WR targets, +0.082 RB carries on the squad denominator; refutation clause not triggered), but ~40-58% of the gap is the season leader's absences, not rotation (A3 failed). Consequence: re-keying to the week estimand beats H-109's squad arm by 0.011 pts MAE…
REFUTED (as a proposed fix). Premise confirmed, consequence reversed: rebasing to the squad denominator is worse on fantasy MAE by 0.016-0.019 pts (Bonf-3 excludes zero in both code states) and on the squad-T split by 0.002-0.004. Do not ship the relabel. What remains open is WHY a wrong-unit target beats the right-unit one -- that is…
TESTED. The board's FACTORS band input is not the unit FACTORS was fitted on: live today 16.2% of carry bands and 25.9% of target bands differ from the target-row fitted unit (moving projections 0.234 [0.192, 0.278] pts/row, RB 0.40), walk-forward 3.5-5.8% in season and 12-24% in weeks 1-3. Fed the served unit walk-forward, fantasy MAE…
TESTED. Pooled nil HELD (+0.009 [-0.005, +0.022]); the age-4+ prediction REFUTED (roster3 +0.093 worse, interval [-0.024, +0.199]). Moving the six l3 columns to an ACT-row unit does not improve the production projection on any stratum, although it improves the raw fp_l3 by 0.54 pts at age 4+. Room totals: projected targets per team-game…
H-173 AS FILED IS REFUTED IN SIZE AND WRONG IN MECHANISM, WITH A NARROWER REAL DEFECT UNDERNEATH. The board matches the published chart QB1 on 30 of 32 teams and follows the chart over last season's starter on 5 of the 7 teams where they differ. The 2 misses are both artefacts of the MIN_LEAGUE_WEEKS=3 hold serving 2025 state into week…
CONFIRMED IN MECHANISM, PREDICTION FAILED ON COUNTS. Three joined tables (rel_coach, regime_gate, kalman_share) have zero 2026 rows and no writer on any live refresh path. At the first board after week 3 five served features lose their source: coach_pos_tgt / coach_pos_car become a literal 0.0, which is BELOW EVERY training value for WR,…
PARTIAL by the registered rule, and the claim's SHAPE is wrong. Direction confirmed and stronger than H-190 said: the MAE-optimal FantasyPros weight at the shipped clamp is 0.95 pooled [0.85, 1.00], QB 1.0, RB 0.95, WR 0.95, TE 0.85; walk-forward per-position weights (0.85-1.0) beat 0.40 by -0.0488 [-0.0700, -0.0227] Bonferroni-8 on…