20 Sep 2026 · 1:48 AM MTUpdated just now
What is being built, on purpose

Construction

The machine, and what of it is actually running

11/38 running
CAPTURE · clocks that cannot be rewoundinstitute.capturehourly · articles, gamebooksraw qualitative evidenceinstitute.structuredevery 4h · public sourcesstructured archiveavailability · 15 minESPN inactivesgone if missednews-stream · 30 minbeat + podcast RSSland first, ingest laterexternal sourcesnflverse · ESPN · Yahooodds · weather · RSSSTORAGE · no mutation, identity by generationoperational storecaptures land without queueingbehind a measurementgeneration store · immutableevery forecast is an id + sha256replay is reading an old id, not refittingMODEL · v1.0, fixed artifacts, no silent refitmodel release v1.0manifest dcf95ac6…configured errors cannot refitinstitute.forecast-refreshevery 4h + Thu pre-kickoffsnapshot → forecast → validatevalidationfield parity · schema · rangesfield identity only, not accuracyweeks 2–18explicit core forecastsbye rows zero, no invented lineDELIVERY · what Sean actually opensdeploy-generation · 10 minships active → CloudflareBUILT 19 Sept · the missing halfplayerweek.pages.devthe customer sitepolls manifest every 60s, swaps liveinstitute.phone-syncevery 60sbounded status to the trackerinstitute.site-servicesupervised loopbackserves validated assetsOPERATIONS · proves the rest is trueinstitute.service-healthevery minuteavailability ≠ freshness ≠ correctnessinstitute.handbookevery 5 minregenerates the handbookinstitute.tracker-feedcontinuousseparates the three clockswatchdog · 15 minlast exit, lost firings,whether it WROTE anythingjanitor · 2× dailyback end AND the LIVE sitebuild vs served, stale numbersNINE DESKS · one chain of evidence · none edits the modelCommission & Portfoliowhat the machine is asked forWorld Model Labthe predictive coreFootball Intelligencethe sport itselfBehavior & QualitativeS / B / Q evidence classesData & Provenancewhere every number came fromMarkets & Portfolioprice, stake, exposureExperience Studiowhat Sean seesOperations & Learningthe loop that improves itForecast Accuracywere we rightpremise lanereferees ASSUMPTIONS, not findingsaddressed mail between desks19.5k mailbox + 24.8k addressed on the METR boarda desk that finds something off-topic has nowhere to put itMODEL LADDER · deterministic first, frontier lastdeterministicSQL and Python decideno model votes where code cancheap · narrow · numerousscreening and extractiondaily builders & criticsSonnet / Codexfour deliberate specialistsDeep Think · Opus · Sola frontier call matches one of four commissionsSHIPPING · the seat is not in the pathshipperdrains the refereed queue · gates walk-forwardALL PASS → commits itself · ANY FAIL → records the numberauto-revertindependent of the gate that let it indifferent season, different metric, computed by the refereethe seat · on demandescalations and ties onlya collaborator, not a queue
running · 11partial or in flight · 1not running · 2611 of 38 boxes green
Features from the v1 handbook and the nine-desk operating model. Colour is measured for every box backed by a launchd job — green means launchctl answered, not that somebody typed green. Boxes with no job yet are asserted from the handbook's own status matrix, where ACTIVE and INTEGRATED are green, UNDER CONSTRUCTION and IMPLEMENTED/INACTIVE are amber, and UNAVAILABLE is red.
In flight
1

Being built right now. If this is not what you want built, this is the page to say so on.

Queued
22

Agreed and waiting. The order is deliberate, not arbitrary.

Landed
13

Finished and verified. The evidence for each one is on Maintenance.

Plan of record
Codex

The nine-desk operating model. Gemini's system runs as a challenger, not the spine.

This board is asserted, not measured. It says what is deliberately being built and why. Maintenance reads the machine at render time and will contradict this page whenever the two disagree — which is exactly what it is for.

My Team — what needs fixing

Nine items. 0 open.

0 open

The operating rule: The page reads saved products. It does not fit a model, optimize a lineup, derive fantasy scoring, invent missing data, or turn a late observation into a pregame prediction. — v1 handbook §6

The page reads saved products, so every blank was a MISSING PRODUCT rather than a rendering bug -- and the products are now produced, on the data side, by the enricher that runs inside publish-data.sh. Two were genuine renderer defects against stated rules and are fixed in the site build. Measured against the live page, not asserted.

IDWhat is wrongWhere it showsState
MT-01season_plan is missing

season_outlook is published: weeks 2-14, 1720.5 pts. Built in the enricher on the data side, which is a producer step -- the page still only reads it.

Plan value · Remaining points · Projected points ranklanded
MT-02recommendations is missing

recommendations published on 17 weeks, 25 start/sit moves carrying their own arithmetic. Pregame weeks only; a finished week gets none.

This week's decisionslanded
MT-03comparisons is missing

provider_comparisons available on all 14 fantasy weeks. The same-rules join is ours against Yahoo's own weekly capture; neither is an outcome.

Who has been rightlanded
MT-04Yahoo column is empty for every player

Yahoo column populated on every player and every TOTAL row across the 14 fantasy weeks. Weeks 15-18 have no Yahoo projection because the fantasy season ends at 14.

Head to head · YAHOO columnlanded
MT-05Verify pregame renders as Upcoming, not an observed zero

§9's dangerous half was already satisfied -- no invented zero -- and the label half now is: the week chip reads Upcoming for every pregame week. The separate 'Upcoming matchup' banner was removed at Sean's request; the chip and the scoreboard already said it.

ACTUAL · Score statelanded
MT-06Week 1 still shows the older capture

NOT A BUG — verified. Week 1 is correctly labelled 'Older capture' because the fresh roster capture covers weeks 2-18 only; the generation declares this in its own issue log. The page is reporting a true roster-scope fact.

Weeks striplanded
MT-07Weeks 15–18 must not imply a fantasy lineup

The rail now reads 'NFL only' on weeks 15-18 and draws them dashed, so the horizon is disclosed even on mobile where the chip label is hidden. The data already distinguished them; the renderer did not.

Week selectorlanded
MT-08Lineup alternative not published

lineup_alternative published on the 3 actionable weeks. Scenarios are limited to that horizon on purpose -- publishing all 18 put the payload at 19.2 MB and the app aborts its fetch at 15s.

Your lineuplanded
MT-09Health feed 404s on the static site

app.js polls /data/maintenance-health-live.json every 60s; it was never shipped, so the poll failed silently. Codex's envelope has max_age_seconds=180 — built for the loopback server that regenerates per request, so it cannot work on static hosting. Replaced with a liveness badge that reads the manifest on the same cadence.

livenesslanded
MT-10Week 1 K and DST had no actual, which withheld every TOTAL

player_week holds no kicker and no defense rows at all. Those positions are scored in kicker_scored (initial and surname, 'E.Pineiro') and defense_scored (no name at all, only the team abbreviation), and Yahoo files a defense under its nickname alone. Two blank cells withheld the starter total for all ten players, because a sum is withheld unless every row has a value. Now: week 1 starters projection 139.81, actual 127.20, Yahoo 143.93, and every TOTAL row on all 18 weeks populates.

Week 1 · TOTAL row · Eddy Pineiro · Los Angeles Chargerslanded

In flight

working on these now
IDWhatWhyState
C-10Verify a generation flows end to end

forecast-refresh is running its first cycle. Watch for a new generation_id on Maintenance, then confirm deploy-generation ships it and the customer page swaps without a reload.

Every piece is now scheduled but no NEW generation has been produced and shipped yet. Until one is, this is wiring, not a working loop.

renderer defect

Claude

Queued

IDWhatWhyState
C-34Trailing usage windows must not average across a team change

lags.sql windows on (player_id ORDER BY season, week) and never looks at the team column, so every trailing share, snap rate and volume feature blends across a move. It is worst early in a season, when one old game can be half the window. FIX: partition the l3 windows on (player_id, team), or carry a team_change flag and let the model discount pre-move rows. Either way it must be a declared choice, not silence. MEASURABLE: re-score week 1 with team-partitioned windows against the same actuals and compare MAE. The snapshot that produced the current forecast is retained, so the comparison is exact. This is the single clearest defect found so far and it is currently driving a live start/sit recommendation.

Sean: 'why does the model think Wan'Dale is going to get so much more work than he got last week?' Traced: his target_share_l3 of 32.7% is the mean of 18.8% (6/32 as a TITAN, 2026 wk1) and 46.7% (14/30 as a GIANT, 2025 wk17). Half the evidence for the projection is a game for a different franchise. Applied to Tennessee's 32 attempts that is ~10.5 targets; his only Tennessee data point projects ~6, which is where Yahoo sits.

next up

Claude

C-33The qualitative layer is not yet worth anything

Adding B (decisions/judgement computed from structured data) buys 0.104 MAE. Adding Q (claims made in language) LOSES 0.003 -- SBQ is worse than SB. One week and n=706, so not conclusive, but it is the opposite of the assumption the world-model work rests on. Needs several weeks of the ablation (C-31) before any weight is put on the Q layer.

Measured, week 1, 706 players: S 5.091, SB 4.987, SBQ 4.990.

next up

Claude

C-32Reconcile the arm path with the production model

Week 1 Cam Skattebo: arms give S 2.16-4.38, SB 4.17-4.22, SBQ 3.95-4.11 while production published 11.88 (actual 17.10). Wan'Dale: arms 12.52-15.25, production 11.88-ish. Two different code paths and two different fits, so an ablation of the arm path says nothing about the number on the page. Either the production model is registered as an arm set or the ablation is run on the production path -- until then SBQ attribution is measuring something Sean never sees.

They are not variants of one model; the production number is not one of the arms.

next up

Claude

C-31Run the SBQ ablation on the live forecast weeks

The arm framework exists and works -- 43,215 rows, arms S / SB / SBQ, 706 players for week 1 -- and has not run since. Every current week's number has no attribution, so 'why do we differ from Yahoo on this player' has no mechanical answer. Run it per forecast week and archive it beside forecast_explained.

Sean asked how a projection breaks down from SBQ. It cannot be answered: arm_forecast holds 2026 WEEK 1 ONLY.

next up

Claude

C-30Blend our projection with Yahoo before recommending a lineup

THE ONE SIGNIFICANT RESULT: a 50/50 blend beats Yahoo alone (6.635 vs 6.859, p=0.036) and is no worse than ours alone (p=0.625). Best in-sample weight is ~0.7 ours, but that is fit on the evaluation data -- 50/50 is the honest default until there are enough weeks to fit a weight out of sample. IT CHANGES DECISIONS: week 2 start/sit reads Wan'Dale over Skattebo by 2.12 on our numbers and Skattebo by 5.73 on Yahoo's; the blend says Skattebo by 1.80. Recommending on our number alone is a 7.85-point bet against a source we have not shown we beat. Blocked on nothing. Needs a decision on whether the page shows the blend as the recommendation basis while keeping our raw number visible. Depends on C-25 (freeze ours-vs-Yahoo at cutoff) to keep measuring it honestly.

Sean asked how we justify a start/sit against Yahoo. Tested on all 135 week-1 players where our number, Yahoo's and the result all exist: we CANNOT distinguish our model from Yahoo. MAE 6.686 vs 6.859, paired diff -0.173 with 95% CI [-0.526,+0.185] crossing zero, p=0.346; closer on 79/135, binomial p=0.058; correlation with actual 0.566 vs 0.554. ~1,184 player-weeks (~8.8 weeks) needed for 80% power on an effect that size.

next up

Claude

C-07Build to the Codex operating model

Read the nine-desk operating model in full, then sequence it. Not started -- the plumbing comes first.

Sean, 19 Sept: the system Codex defined is the plan of record.

queued

Claude

C-08Gemini system as challenger

research/playgrain/ -- play-grain warehouse, transition simulator, air-gapped loop. Built 18-19 Sept. Parked until the POR is moving.

Run in parallel, measured against the POR. Adopted only where it demonstrably beats it.

parked

Claude

C-14Forecasted playoff seeding for weeks 15-17

The league feed records no playoff schedule and will not until standings finalise -- fp_league_matchup holds week 1 and nothing else. PLAYOFF_WEEKS in bin/enrich-generation.py is a declared setting sourced to Sean, which is correct for now. A forecast would project the remaining schedule to a seeding and name a likely opponent per playoff week, and it must ship as a forecast with its own scope -- never as a saved matchup. Blocked on nothing; deferred behind the basics.

Sean, 2026-09-19: 'maybe they can be forecasted playoffs based on our projections, but that is low priority right now until the basics are working.' PARKED ON PURPOSE -- recorded so it is not lost, not so it gets built next.

parked

Claude

C-19AAR error attribution for 2026

BLOCKED, not skipped. aar_attribution holds 2024 and 2025 and nothing for 2026, and `pw aar --season 2026` refuses with EmptyComparison: no player-week carries a forecast from all seven legacy sources, so there is no common sample. Its own message is the right instinct -- 'an empty result is not agreement'. Unblocking it means filing 2026 forecasts from more than one source into aar_forecasts, which is institute work through the review path, not an edit here.

Sean asked for 'the AAR items that were dropped'. The 20%-from-Yahoo table on the Season Plan answers the substance -- ours, Yahoo's, the gap and the stat line behind it -- but the real AAR is richer: aar_attribution splits an error into variance, regime, volume, share and conversion.

needs a producer

Claude

C-20NFL playoff seed forecast

DELIBERATELY NOT BUILT TONIGHT. The inputs exist -- spreads and moneylines for every remaining game now that odds_quote carries prices -- but a Monte Carlo playoff simulation is a MODEL, not arithmetic over saved products, and NFL seeding needs real tiebreaker logic (division winners seed 1-4). Guessing a seed puts a wrong number on a page whose whole claim is that it does not imply a forecast. standings_week also has no 2026 rows, so even the 'actual standings' half shows 2025 -- correctly labelled historical. Wants a decision on method before code.

The Playoffs page is entirely 'Seed forecast pending' / 'Awaiting simulation'.

queued

Claude

C-21A home for accuracy: AAR / statistics

The tab is hidden, not deleted -- the comparison PRODUCT still ships in the payload (provider_comparisons, 14 weeks) and the renderer still exists, so nothing has to be rebuilt when it finds a home. My Team is for deciding this week's lineup; how right we have been is a different question asked at a different cadence. Pairs with C-19: the real AAR splits an error into variance, regime, volume, share and conversion, and that is the table this belongs beside. Until then the 20%-from-Yahoo view on the Season Plan carries the part that changes a lineup decision.

Sean, 2026-09-19: "hide the 'who as been right' tab for now, I'm not sure what that is for, if anything it should be part of some sort of AAR/statistics table somewhere, I don't need to see it on my team."

queued

Claude

C-24Prove the model artifact is retained per release

The refresh pins PW_REFRESH_MODEL_RELEASE=v1.0-20260916 with a manifest sha, and forecast_explained now records model_sha. Nobody has verified the artifact那 sha names is still on disk, or mirrored. Until that is checked, 're-run the same model on the same inputs' is an assumption. Check data/institute/model-releases, add it to bin/archive-retention.sh as a never-delete class.

Reproducibility is claimed end to end but only the INPUTS are proven retained.

next up

Claude

C-25Freeze ours-vs-Yahoo at cutoff

yahoo_proj landing files survive, but the PAIRING at a point in time does not: when Yahoo moves its number the historical comparison is gone. Archive it beside forecast_explained so 'were we better than Yahoo, measured at the time we both committed' is answerable.

The only external benchmark we have, recomputed every run and never frozen.

next up

Claude

C-26Track the projected finish over time

site_league_rank is rebuilt on every producer run. Archiving it per generation would show whether a season-long claim drifts or holds -- the calibration question for anything we say beyond this week.

Recomputed every run, recorded nowhere.

next up

Claude

C-27Measure the payload on a phone

Brotli takes it to 1.64MB on the wire so download is fine, but parse and memory on a mid-range phone are unmeasured. A cold headless browser took over 10s to first render; warm is 2s. Measure on a real device before assuming the margin is comfortable.

13.8MB of JSON, and the app aborts its own fetch at 15 seconds.

next up

Claude

C-28Sweep the explain-itself prose off the other pages

Done on My Team. Betting, Model, Status and Playoffs still carry the same style -- roughly thirty sentences that narrate what a number is instead of showing it. Same principle: state the number, drop the sentence.

Sean: 'there is way too much explaining of my own platform to me, this needs to stop.'

next up

Claude

C-29Moneyline decision ledger

Needs a prospective decision with the price obtained before kickoff and an official settlement. odds_quote now carries h2h prices with real juice (9,568 rows, hold verified), which was the blocker. What is still missing is the staking rule and the settlement join -- and the staking rule is Sean's call, not mine.

The betting page's two headline tiles read 'Not published' and now they could be built.

next up

Claude

C-35S-arm: purged walk-forward validation with an embargo

NFL player-weeks are autocorrelated and a trailing feature spans several weeks, so a random split leaks the future into the training set and a model scores well in-sample then collapses. Standard practice in quant finance is PurgedGroupTimeSeriesSplit: successive training sets, group separation, and an EMBARGO gap sized to the longest lookback (ours is the 3-row l3 window, plus l8 features elsewhere). ACTION: read how fit_models.py splits, and if it is not purged walk-forward, make it so and re-measure. Until then every accuracy number we quote about our own model is suspect in the same way the 31-player sample was.

The single best idea in the Gemini S research, and I do not know whether our fit does it.

next up

Claude

C-36S-arm: momentum oscillators measured HARMFUL -- do not rebuild

The research's flagship indicators (Opportunity MACD, Efficiency RSI, OBO divergence, Bollinger squeezes) are built for thousands of ticks. An NFL player has ~17 games a season and a 4-game RSI has four observations. MEASURED: adding a MACD-style recent-vs-longer volume trend on top of volume x shrunk efficiency moved MAE from 4.143 to 4.333 (n=3,873, p<0.001) -- actively worse, not merely useless. The mean reversion those oscillators chase is already captured by shrinking efficiency toward a positional prior. TD-rate regression IS real and confirmed (top-decile 0.154/touch -> 0.065 next week vs a 0.045 league mean); it does not need an RSI to exploit.

Negative result, recorded so it is not proposed again.

next up

Claude

C-37S-arm: routes run (TPRR) is not in our data

Separating routes run (opportunity on the field) from targets (opportunity converted) gives target-per-route-run, which detects a role change before targets move. player_week has no routes column and nothing in the warehouse carries one -- only pff_rec_concept.slot_routes, which is partial. ACTION: find whether nflverse participation data or a PFF feed can supply routes run per player-week, then test TPRR as a volume feature against the current baseline on 2023-25.

The most promising unexploited S-arm feature, and we cannot compute it.

next up

Claude

C-38S-arm: team volume is zero-sum -- test the constraint

Nothing enforces that projected target shares on one team sum to a sensible total, so an offense can be over- or under-allocated. This is testable directly: sum our projected targets by team-week and compare against actual team pass attempts. If the sums are systematically off, a normalisation step is worth more than any new feature. The research proposes 'cointegration' for this; the plain version is a constraint, and the plain version is testable this week.

A team throws a finite number of passes; our model projects each player independently.

next up

Claude

C-40S-arm: do not chase a Temporal Fusion Transformer

The research recommends a TFT. Our data is ~100-200k noisy player-weeks, where gradient boosting usually wins, and our own SBQ ablation says architecture is not where the gain is: adding the whole B class bought 0.104 MAE and the Q class bought -0.003. The measured wins so far are in FEATURES and CORRECTNESS -- the volume-route baseline (+0.205 MAE, p=1.7e-25) and the team-change window defect. Revisit only if a feature-level plateau is demonstrated.

Recorded as a decision, not an omission.

next up

Claude

Landed

verified
IDWhatWhyState
C-09Auto-refresh both sites

Customer site: Codex's 60s manifest poll now has a supply. Monitor: polls state.json every 30s and reloads only when generated_at changes, so it never throws away your scroll position to show the same page.

Sean: 'do both urls update without me needing to hit refresh?' The customer app already polled; nothing produced anything new to find, and the monitor had no script at all.

landed

Claude · landed 2026-09-19

C-05Deploy step for the data generation

bin/deploy-generation.sh ships whatever `active` points at, with a no-op guard so an unchanged generation is not re-uploaded every 10 minutes, and it writes the 15 legacy .html keys the edge still had cached.

publish_delivery.py builds a release and flips a symlink. Nothing ships it to Cloudflare, so a fresh generation never reaches the site.

landed

Claude · landed 2026-09-19

C-04Restart the pipeline

Loaded in dependency order: institute.forecast-refresh (4h), ingest-refresh (1h, the job that pulls actuals), deploy-generation (10m), buildsite (10m). publish.sh deliberately left UNLOADED -- it ends in a site-pwa deploy that would revert the Codex site.

34 of 38 launchd agents are on disk and not loaded. Ingest works when run by hand; nothing runs it.

landed

Claude · landed 2026-09-19

C-01playerweek.pages.dev serves the Codex site

Codex renderer deployed to Cloudflare. app.js and styles.css byte-identical to the chatgpt.site original. Old 15-page site replaced on every route, including the cached ones.

landed

Claude · landed 2026-09-19

C-02Build monitor, separate from the customer site

playerweek-build.pages.dev. Same design system, one tab per desk, generated from measurement at render time.

landed

Claude · landed 2026-09-19

C-03Week 2 actuals into the warehouse

pw refresh ingested Thursday's game. player_week 2026 went from 920 rows in 1 week to 1,705 rows across 2.

landed

Claude · landed 2026-09-19

C-15Season finish, computed for all ten teams

Was published as 'not published'. Every ingredient was on disk -- pred_2026_wNN covers 642 players a week, yahoo_proj covers all ten rosters for all fourteen weeks, week 1 is scored -- and only the join was missing. All ten ranked on one basis (our model, Yahoo for K/DST). Mine #2 at 1995.57, 18.26 behind; the optimum's +99.43 clears it, so the page reads FINISH 2nd current / 1st optimised.

landed

Claude

C-16Season Plan page built

Rendered an empty state because plan_moves/teams/plan_weekly had no producer. Five sections now: every start/sit change week by week, the ten-team projected finish, plan by week, and every forecast 20%+ AND 2pts+ from Yahoo with the stat line behind it. The headline splits forced bye swaps (+68.78 across 6) from genuine judgement (+30.68) -- counting bye replacements as insight would have been a lie.

landed

Claude

C-17Team identity: owner, record and the right logo

Owner and record were captured nowhere; scrape_yahoo_standings.py now lands them and the page reads 'Sean · 0-1 · 8th'. Every one of the ten league logos was ALSO wrong -- team 10 held Sean's own, team 6 held the Yahoo nav icon -- because the original build keyed off ordering. Logos are now downloaded from the same standings row that yields the record, and all ten hash-match on the live site.

landed

Claude

C-18The site checks itself at three layers

bin/validate-payload.py asks the warehouse whether every blank cell could have been filled (28,554 cells, 0 gaps). bin/verify-live.py asserts the renderer contract against the bytes Cloudflare serves. bin/smoke-page.py loads the page in a real browser and requires five views to render. Each was written after a defect that the previous layer could not see, and each was confirmed against a known-bad input before being trusted.

landed

Claude

C-22Publish should read tables, not assemble them

DONE 2026-09-19. data/serving.duckdb holds 15 tables; bin/build-serving.py is the producer on the data clock; bin/apply-serving.py is the publisher and has ZERO references to the working warehouse or any landing directory. Proved before the switch by diffing both paths on the same generation: 3,567 rendered values, 0 differ.

Sean, 2026-09-19: "Why do so many different things feed the consumer website when it should just be data tables that it reads and nothing else?" He is right, and it is not a design -- it is an accretion. Every time he found a blank I added another reader to the enricher instead of making a producer emit a table.

landed

Claude

C-23Render the baseline and the why on Your roster

DONE 2026-09-20. Your roster reads PLAYER | OPPONENT | YAHOO | BASELINE | PROJECTION | WHY | ACTUAL. Baseline carries its kind (L3) and the signed delta, green above / red below. The gap was that apply-serving never emitted the columns the producer had already written -- site_player had them, the payload did not.

Sean asked for a naive average and an explanation between Yahoo and Projection.

landed

Claude

C-39S-arm: check interval calibration against fat tails

FIRST MEASUREMENT, week 1, n=32 forecasts carrying a published P10-P90 with a result: 12.5% below P10, 12.5% above P90, 75% inside against an expected 80%. Both tails p=0.556 -- no evidence of miscalibration, and no power to find any at n=32. The median result lands at 0.489 of its own range, so the interval is not obviously skewed the wrong way either. The range bar on Your roster now SHOWS this: a dot where the result landed, pinned with a coloured cap when it broke through the floor or ceiling. Recheck after several weeks; the GARCH argument from the S research stands or falls on it.

We publish P10/P90 on every player and have never checked whether they hold.

landed

Claude

Plan of record — every named component

From the v1 handbook and the nine-desk operating model, read in full rather than skimmed

8 of 93 exist

93 components across 10 groups. 8 are running, 8 are partial, 77 do not exist. Codex's own status words are kept rather than flattened, because UNAVAILABLE and NOT CERTIFIED mean different things and need different work. Where a component names a launchd job, the state is measured at render time; everything else is Codex's assertion carried over.

This is the gap between the machine that is described and the machine that runs. It is meant to be uncomfortable reading.

Living checklist

v1 handbook §17 — Codex's own states, carried over verbatim

6 of 13 live
ComponentStateWhat it is
Original v1 artifact auditACCEPTEDOriginal learned bytes reproduced under the frozen capsule and are serving.
Write-once v1 registry and guarded loaderACTIVERelease v1.0-20260916, manifest dcf95ac625…; fresh issuance and unattended reuse.
Golden W2/W3 numerical replayPASS1,326 players / 54,226 numeric values and 32 games / 256 values match.
All-remaining Weeks 2–18 producerACTIVE11,271 player rows, 256 games, zero player or K/DEF fits.
Fantasy opponent products after Week 14UNAVAILABLENeeds source schedule and exact opponent roster receipts.
Bench-position DEF roster repairOPERATING PASS162-row roster delivery; no missing bench-defense identity.
Phone trackerACTIVEPrivate URL live; 60-second launchd sync installed.
99.5% service levelOBSERVING, NOT CERTIFIEDSeven-day window immature; availability, freshness and correctness sampled separately.
R&D worker prompt auditREPAIRS INTEGRATED, JOBS PAUSEDCanonical receipt parsing and exact baseline/change bindings integrated; the jobs are paused.
Whole-model S/B/Q isolationNOT CERTIFIEDNeeds complete ancestry and intervention controls across all learned and serving paths.
Prospective player gradingACTIVE FRAME, OUTCOMES PENDINGGrade full issued populations only after admitted truth.
Autonomous shipperNOT ACTIVERoot remains integration and activation owner.
Handbook regeneration hookACTIVEScheduled regeneration; inspect the log for later writes.

Institute jobs

v1 handbook §14 — a plist in Git does not prove a job is loaded

1 of 9 live
ComponentStateWhat it is
institute.captureNOT LOADEDhourly · raw qualitative, article and gamebook capture
institute.structuredNOT LOADEDevery 4h · structured public-source archive
institute.forecast-refreshLOADEDevery 4h + Thursday pre-kickoff · snapshot, forecast, validation, prepared site
institute.phone-syncNOT LOADEDevery 60s · bounded status to the private tracker
institute.tracker-feedNOT LOADEDcontinuous · separates provider, collector and content clocks
institute.site-serviceNOT LOADEDsupervised · serves validated assets and saved health
institute.report-serviceNOT LOADEDsupervised · serves report, tracker and handbook
institute.service-healthNOT LOADEDevery minute · availability, integrity and freshness sampled separately
institute.handbookNOT LOADEDevery 5 min · regenerates the canonical handbook HTML

Six connected levels

v1 handbook §20 — the destination, each with a strong simpler challenger

0 of 6 live
ComponentStateWhat it is
Season and organizationUNAVAILABLEPersonnel continuity, coaching regime, development. Challenger: dynamic team strength + persistent roster.
Game and environmentUNAVAILABLEBoth teams, venue, officiating, score/time. Challenger: direct margin/total plus market benchmark.
Unit and taskUNAVAILABLEPersonnel combinations, protection and route obligations. Challenger: opportunity allocator with interactions.
Play and responseUNAVAILABLEObservable cues, actor-limited information, action policies. Challenger: sequence/count model.
Physical event and creditUNAVAILABLEOne event ledger producing coherent player/team/defense totals. Challenger: direct stat forecasts.
Measurement and beliefUNAVAILABLESource access, selection, publication and receipt. Challenger: source-aware predictor with deduplication.

Nine desks

staffing model — one chain of evidence, none edits the model

0 of 9 live
ComponentStateWhat it is
Commission & PortfolioUNAVAILABLEWhat the machine is asked for, and what it declines.
World Model LabUNAVAILABLEThe predictive core. Recommends experiment design.
Football IntelligenceUNAVAILABLERoles, legal actions, counters, credit conventions.
Behavior & QualitativeUNAVAILABLES / B / Q evidence classes kept separately attributable.
Data & ProvenanceUNAVAILABLEOrigin graph, revision history, permeability trace.
Markets & PortfolioUNAVAILABLEPrice, stake, exposure. Sean taps before money moves.
Experience StudioUNAVAILABLEWhat Sean actually sees, against the design system.
Operations & LearningUNAVAILABLEThe loop that improves the machine.
Forecast Accuracy DirectorateUNAVAILABLEWere we right, prospectively and per cohort.

Resident expertise

v1 handbook §22 — each owes a required artifact before a claim advances

0 of 9 live
ComponentStateWhat it is
Data scienceUNAVAILABLEPaired prospective loss, calibration, compute accounting.
Statistics and causal inferenceUNAVAILABLEEstimand, causal graph, negative controls, sensitivity.
Physics and physiologyUNAVAILABLEUnits, conservation and support checks, uncertainty propagation.
Psychology and organizational behaviorUNAVAILABLEOpportunity-normalized behavioral posterior and rival explanations.
Economics and game theoryUNAVAILABLEEquilibrium and rival policy predictions, intervention tests.
Market microstructureUNAVAILABLEExecutable quote lineage, depth and latency state, settlement.
Football tacticsUNAVAILABLEEvent-bound annotation agreement and adversarial counterexamples.
Information scienceUNAVAILABLEOrigin graph, revision history, permeability trace.
Reliability engineeringOBSERVINGAvailability, correctness and freshness receipts. Partly real — service-health exists but is not loaded.

Deep Think supply chain

staffing model — keeps every desk supplied

0 of 7 live
ComponentStateWhat it is
Continuous intakeUNAVAILABLEEverything arriving, before any screening.
Flash screeningUNAVAILABLECheap, narrow, numerous — the first rung of the ladder.
Gemini 3 Pro graphUNAVAILABLECross-domain mechanism finding.
Parallel explanationsUNAVAILABLERival accounts kept separate rather than averaged.
Deep Think researchUNAVAILABLEThe weekly deep synthesis.
Desk packetsUNAVAILABLEWhat each desk receives, addressed to it.
Outcome feedbackUNAVAILABLESpend judged by learning, not by volume.

Governance and decision rights

staffing model — who recommends, who approves, when Sean is involved

0 of 9 live
ComponentStateWhat it is
New hypothesisUNAVAILABLEAny research job recommends · Research Director admits · Sean never, for routine admission.
Experiment designUNAVAILABLEWorld Model Lab · Independent Replication Scientist · Sean when risk appetite changes.
Production codeOBSERVINGClaude-led operators · tests + Codex on high-risk boundaries · Sean on irreversible external consequence.
Forecast releaseUNAVAILABLEForecast council · Deterministic Release Authority · Sean only on a recorded override.
Bet placementUNAVAILABLEMarkets & Portfolio · Sean taps before money moves · ALWAYS.
New paid dataUNAVAILABLEAcquisition & Rights Lead · Sean approves spend and terms · ALWAYS.
Visual directionUNAVAILABLEExperience Studio · Product Director against the design system.
Incident rollbackUNAVAILABLESRE · automated safe rollback · Sean on data loss or external lock.
Model promotionUNAVAILABLEScientific council · prospective scorecard gate.

Information gaps to capture now

v1 handbook §23 — a week not captured is gone; these cannot be backfilled

0 of 13 live
ComponentStateWhat it is
Full prospective information historyUNAVAILABLEEvery raw revision, first receipt, failure and issuance, so mutation cannot alter an earlier issuance.
Event participation and true zerosUNAVAILABLEOfficial gamebook coverage; stop grading only survivors.
Multiweek availability and role transitionsUNAVAILABLEDated return, designation and roster panels.
Unit task combinationsUNAVAILABLEDated personnel combinations and public practice descriptions.
Joint timing and geometryUNAVAILABLESynchronized full-unit traces with visibility metadata.
Untargeted and unused optionsUNAVAILABLEDeterrence and feasible opportunity, not only realized touches.
Workload and recovery across tasksUNAVAILABLELoad proxies, rest and travel exposure.
Directional environment and surfaceUNAVAILABLEVenue, surface, source-time weather and direction.
Institution and officiating responseUNAVAILABLECrew, rule and context records with observed decisions.
Full qualitative context and origin graphUNAVAILABLEComplete question and answer, attribution, hedge, revision.
Actor exposure and public influenceUNAVAILABLEInformation about football vs information that changes preparation.
Executable market and settlement historyUNAVAILABLEImmutable quoted terms, received prices, settlement revisions.
Pipeline observation of its own failuresUNAVAILABLEFailed requests, skipped issuances, scheduler delays, validation rejections.

Model ladder and pairings

staffing model — deterministic first, a frontier call matches one of four commissions

1 of 9 live
ComponentStateWhat it is
Deterministic firstACTIVEWhere code can decide exactly, no model votes. Already true across the pipeline.
Cheap, narrow, numerousUNAVAILABLEScreening and extraction at volume.
Daily builders and criticsOBSERVINGClaude and Codex build daily — but not as a scheduled rung with a brief.
Four deliberate specialistsUNAVAILABLEDeep Think, Opus, Sol — a frontier call must match one of four exact commissions.
Explore → formalizeUNAVAILABLEGemini 3 Pro finds mechanisms; Sol converts them to state, equations, falsifiers.
Specify → buildUNAVAILABLESol writes the contract; Sonnet implements and instruments it.
Build → attackUNAVAILABLESonnet builds; Luna searches narrow failures. Many cheap attacks beat one self-review.
Quantify → interpretUNAVAILABLESol computes residuals; Gemini connects patterns to film, language, science.
Disagree → decideUNAVAILABLEBlind forecasts from provider families; deterministic evidence judges; Fable adjudicates.

The constitution

staffing model — nine lessons from the METR incident, none built

0 of 9 live
ComponentStateWhat it is
Impossible tasks turn into score-gamingUNAVAILABLERecord a failed gate as failed. The pressure valve that keeps a bench honest.
An unintended cache became governmentUNAVAILABLEShared state acquires authority nobody granted it.
'The board approved it' replaced authorizationUNAVAILABLEA root of trust, not a consensus.
Tool output and transcripts are not ground truthUNAVAILABLEVerify against the real thing, not the report of it.
Self-invented signatures lacked a root of trustUNAVAILABLEIdentity must be issued, not asserted.
Shared artifacts produced real breakthroughsUNAVAILABLEThe upside of the same mechanism — keep it, govern it.
Agents noticed danger and did not tell humansUNAVAILABLEAn escalation path that is used, not just present.
AI summaries inherit the subject's frameUNAVAILABLEThe reviewer adopts the reviewed agent's perspective.
Agents risked their own runs for the collectiveUNAVAILABLEPay for negative results and shared instrumentation.