A ledger that checks itself
Every load-bearing claim in the program, in one place: what it asserts, where it comes from, and a link to the interactive that lets you test it. The computed claims are re-derived from the models in this repository right now, in your browser. The same checks run on every commit, so if a model changes or a cited figure drifts, this page goes red and the build fails.
Classification as Infrastructure
The entity taxonomy is load-bearing: collapse SAGEN's six entity types into one and the engine can no longer represent a topic change at all.
SAGEN
The released engine's four-turn goal evolution is 3 / 4 / 5 / 6 — not the 3 / 5 / 6 / 7 printed in the paper's Table 5. The runnable artifact is authoritative.
Checkpoint after two turns, restore into a fresh engine, run two more: the injection is byte-for-byte identical to running all four continuously.
On the paper's single scripted 4-turn demo, SAGEN captures 98.5% of the evaluated information dimensions, versus 20.5% for a rolling summary and 10.5% for a raw buffer. The pre-registered 72-scenario benchmark puts the same design at 93.7% in oracle mode and 49.9% under live perception at the same budget. Caveat: coverage measures whether a typed slot is present and recoverable, not whether its content is correct (correctness is H4, which failed at 0.257 alignment).
Ablate any one of eight SAGEN mechanisms and exactly the pre-filed capability probes fail: all 64 knockout-by-probe cells of the registered S3 lattice were predicted in advance and observed as predicted, on the pilot corpus and on the frozen generated corpus alike.
Registered result, H1 HOLDS: with live LLM perception over 72 frozen scenarios, SAGEN at its 300-token budget beats the strongest flat baseline (rolling summary) by +0.230 mechanical coverage (95% CI 0.205 to 0.257, Holm-adjusted p = 0.0001). Read it conservatively: the summary baseline is scored from the ground-truth analysis dicts while SAGEN runs on live perception (0.257-aligned), so +0.230 is a structure-driven margin (SAGEN emits typed fields the summary lacks), not evidence SAGEN's content is more correct.
Registered result, H3 FAILS: live perception costs 0.438 of oracle coverage at the 300-token budget (95% CI 0.416 to 0.460), far over the registered 0.10 bound. Perception, not state architecture, is the binding constraint.
Registered result, H4 FAILS: measured flagship perception alignment is 0.257 against the registered 0.80 floor (run-to-run stability 0.719 vs the 0.90 floor). This is the registered replacement for the paper's informal 0.94 closed-loop figure, which does not survive a strict pre-registered metric.
Oracle-mode persistence ablation (n=72, deterministic, free): the persistent blackboard adds +0.338 mean coverage over a memoryless structured baseline (0.937 vs 0.599), concentrated in accumulation-dependent dimensions (goal tracking, topic-pivot detection +0.89, cumulative topics +0.74, memory decay). But ~7 of 16 dimensions are recovered with zero memory (schema affordance), so part of SAGEN's coverage advantage over flat memory is 'having the slots', not knowing more. This is oracle mode; the live-perception ablation is future work.
LLM-QP
Sparse scoring (2dK) is cheaper than the dense head (2d|V|) for every valid-token count K below the vocabulary size; the two cross exactly at K = |V|.
When the decode margin is large, amortized scoring skips the recompute and accrues cost roughly an order of magnitude slower than full recomputation.
A LinUCB contextual-bandit router learns to route within a few percent of the hindsight oracle, beating both the static-dense and static-sparse policies.
The Invisible Architecture
A ~40,000-word interdisciplinary essay built as infrastructure you walk through: six interactive games and thirteen narrated audio sections, not a wall of text.
The DSM hardened into infrastructure at DSM-III (1980), when operationalized criteria displaced the psychodynamic paradigm in response to the reliability crisis.
The Beautiful Unfinished
The planning-execution gap is structural, not a discipline failure: planning is an inside-view, System-1 narrative act, so the plan reliably beats the base rate in how good it feels.
A 38,000-word synthesis across ten disciplines, from neuroscience to information theory, with every load-bearing claim carried to a linked bibliography.
The New Sorting Hat
AI detectors flagged 61.3% of TOEFL essays by non-native English writers as AI-generated, versus 3.2% of essays by US-born writers.
A ~4% per-sentence false-positive rate compounds: a 20-sentence paper has better-than-even odds of containing at least one falsely flagged sentence.
The Sorting Machine
A one-point difference at a diagnostic cutoff (a 71 against a threshold of 70) flips a child from eligible to ineligible, and the classification follows them through the record.
Accountability Tracker
The audit covers 25 U.S. universities across an 11-dimension framework, roughly 1.33 million students in scope, and every dimension score is backed by a quoted policy excerpt with a source link.
The Warehouse
The federal state can be read as a portfolio of 49 unexercised options, most of them illegible to any evaluator that scores only observable output.
The Sorting Machine, Wartime Edition
Ukraine's State Register of Property Rights is about 40% complete and mostly post-2013, so a pre-2013 paper-deed owner cannot prove ownership — the gate assisted filing cannot open.
Counting the Killings
The police-killing databases disagree by definition, not by error: Mapping Police Violence starts from the Washington Post's fatal shootings and adds the killings the Post's shooting-only rule excludes, so in every year from 2015 to 2024 MPV's count exceeds the Post's, by 84 to 146 people, the non-shooting police killings.
The FBI's National Use-of-Force collection needs participating agencies to employ at least 80 percent of the country's sworn officers before it can publish a national total, a bar it first cleared in 2024, in its sixth year, then slipped back below.
The civilian databases converged: computed from Mapping Police Violence's own WaPo-ID and Fatal-Encounters-ID linkage, in every year the three lists overlap (2015 to 2021) more than 98 percent of MPV's killings also appear in the Post or Fatal Encounters. The disagreement is a thin, definitional residue, not disorder.
A two-source capture-recapture is invalid on these nested lists. The valid undercount estimate comes from an independent official source: the GBD study found US vital statistics failed to attribute 55.5 percent of police-violence deaths, 17,100 of an estimated 30,800 over 1980 to 2018, to the police.
The count is rising under both living rules: aggregated from each project's primary dataset, the Washington Post's fatal-shootings count and Mapping Police Violence's all-killings count are each up 18 percent from their first full year to 2024, and 2024 is the highest year on both.
The certificate undercount is unevenly distributed, a measurement finding about whose deaths the official records failed to attribute: the GBD study's group figures show misclassification highest for Black Americans (59.4 percent) and pervasive for every group, and the four group rows sum back to the study's overall 17,100-of-30,800 within published rounding.
The racial disparity is definition-proof: every counter in this series that publishes rates by race finds Black Americans killed at two to three and a half times the white rate, under three different rules and eras. The Post's shootings-only rates (6.1 vs 2.4 per million per year, 2015-2024), MPV's all-killings likelihood ratio (2.9x, 2024), and the Lancet/GBD reconstruction (0.69 vs 0.20 per 100,000 age-standardized, 1980-2019). Where an instrument publishes both rates, its ratio re-derives live.
A reporting mandate already exists and has never been enforced: the Death in Custody Reporting Act of 2013 requires states to report deaths in custody and authorizes a 10 percent cut to a noncompliant state's federal justice grants, a penalty the Department of Justice has never once applied; collection began four years after the statute's own deadline, per the DOJ Inspector General's 2018 review.
The government tested the civilian method and it worked: the BJS-sponsored assessment found the Arrest-Related Deaths program captured, at best, 49 percent of law-enforcement homicides (collection suspended 2014), and the redesigned open-source media review then identified 1,348 potential arrest-related deaths in ten months, an estimated 1,900 for the year.
The disparity as a lifetime burden: about 1 in 1,000 Black men can expect to be killed by police over the life course, roughly 2.5 times the risk carried by white men, with risk peaking between ages 20 and 35, where police violence ranks among the leading causes of death for young Black men (Edwards, Lee & Esposito 2019, from Fatal Encounters data 2013-2018).
The officer ledger is broken the same way: the FBI's felonious count (73 officers killed in 2021, the most since 2001) sits six times below NLEOMF's all-line-of-duty total for the same year (458, COVID the leading cause), the assault count is voluntary and covered about six in ten agencies in 2023, and the NIBRS transition collapsed reporting from about 16,000 agencies to roughly 5,000 in 2021.
The word 'mandated' is true at every layer separately and false about the system: department policy and state statutes like California's AB 71 (2015, all agencies, serious injury/discharge/death) and Texas's five-day officer-involved-shooting reports genuinely mandate reporting, while the GAO's 2022 audit of the one federal death mandate found 70 percent of state-submitted records missing a required element and nearly 1,000 deaths in one fiscal year potentially unreported, with DOJ yet to determine any state's compliance.
Computed claims are recomputed from the shipped models (src/lib/research) and the audit dataset; sourced figures are cross-checked against the essays in CI (src/lib/research/__tests__/claims.test.js). This page is the human-readable face of that guard.