Research · The Claims

A ledger that checks itself

Every load-bearing claim in the program, in one place: what it asserts, where it comes from, and a link to the interactive that lets you test it. The computed claims are re-derived from the models in this repository right now, in your browser. The same checks run on every commit, so if a model changes or a cited figure drifts, this page goes red and the build fails.

34 claims0 / 13 re-verified liveall guarded in CI
01

Classification as Infrastructure

2 pivots → 0rechecking…

The entity taxonomy is load-bearing: collapse SAGEN's six entity types into one and the engine can no longer represent a topic change at all.

02

SAGEN

goals 3 / 4 / 5 / 6rechecking…

The released engine's four-turn goal evolution is 3 / 4 / 5 / 6 — not the 3 / 5 / 6 / 7 printed in the paper's Table 5. The runnable artifact is authoritative.

byte-identicalrechecking…

Checkpoint after two turns, restore into a fresh engine, run two more: the injection is byte-for-byte identical to running all four continuously.

98.5% vs 20.5% vs 10.5%cited

On the paper's single scripted 4-turn demo, SAGEN captures 98.5% of the evaluated information dimensions, versus 20.5% for a rolling summary and 10.5% for a raw buffer. The pre-registered 72-scenario benchmark puts the same design at 93.7% in oracle mode and 49.9% under live perception at the same budget. Caveat: coverage measures whether a typed slot is present and recoverable, not whether its content is correct (correctness is H4, which failed at 0.257 alignment).

64/64 predictedrechecking…

Ablate any one of eight SAGEN mechanisms and exactly the pre-filed capability probes fail: all 64 knockout-by-probe cells of the registered S3 lattice were predicted in advance and observed as predicted, on the pilot corpus and on the frozen generated corpus alike.

+0.230 live advantage✓ recomputed in CI

Registered result, H1 HOLDS: with live LLM perception over 72 frozen scenarios, SAGEN at its 300-token budget beats the strongest flat baseline (rolling summary) by +0.230 mechanical coverage (95% CI 0.205 to 0.257, Holm-adjusted p = 0.0001). Read it conservatively: the summary baseline is scored from the ground-truth analysis dicts while SAGEN runs on live perception (0.257-aligned), so +0.230 is a structure-driven margin (SAGEN emits typed fields the summary lacks), not evidence SAGEN's content is more correct.

tax 0.438 vs 0.10 bound✓ recomputed in CI

Registered result, H3 FAILS: live perception costs 0.438 of oracle coverage at the 300-token budget (95% CI 0.416 to 0.460), far over the registered 0.10 bound. Perception, not state architecture, is the binding constraint.

0.257 measured (0.94 printed)✓ recomputed in CI

Registered result, H4 FAILS: measured flagship perception alignment is 0.257 against the registered 0.80 floor (run-to-run stability 0.719 vs the 0.90 floor). This is the registered replacement for the paper's informal 0.94 closed-loop figure, which does not survive a strict pre-registered metric.

+0.338 persistence deltacited

Oracle-mode persistence ablation (n=72, deterministic, free): the persistent blackboard adds +0.338 mean coverage over a memoryless structured baseline (0.937 vs 0.599), concentrated in accumulation-dependent dimensions (goal tracking, topic-pivot detection +0.89, cumulative topics +0.74, memory decay). But ~7 of 16 dimensions are recovered with zero memory (schema affordance), so part of SAGEN's coverage advantage over flat memory is 'having the slots', not knowing more. This is oracle mode; the live-perception ablation is future work.

03

LLM-QP

crossover at K = |V|rechecking…

Sparse scoring (2dK) is cheaper than the dense head (2d|V|) for every valid-token count K below the vocabulary size; the two cross exactly at K = |V|.

~10x cheaperrechecking…

When the decode margin is large, amortized scoring skips the recompute and accrues cost roughly an order of magnitude slower than full recomputation.

near-oracle (≈1.04x)rechecking…

A LinUCB contextual-bandit router learns to route within a few percent of the hindsight oracle, beating both the static-dense and static-sparse policies.

05

The Invisible Architecture

6 games · 13 audiocited

A ~40,000-word interdisciplinary essay built as infrastructure you walk through: six interactive games and thirteen narrated audio sections, not a wall of text.

DSM-III, 1980cited

The DSM hardened into infrastructure at DSM-III (1980), when operationalized criteria displaced the psychodynamic paradigm in response to the reliability crisis.

06

The Beautiful Unfinished

inside vs outside viewcited

The planning-execution gap is structural, not a discipline failure: planning is an inside-view, System-1 narrative act, so the plan reliably beats the base rate in how good it feels.

38,000 words · 10 fieldscited

A 38,000-word synthesis across ten disciplines, from neuroscience to information theory, with every load-bearing claim carried to a linked bibliography.

07

The New Sorting Hat

61.3% vs 3.2%cited

AI detectors flagged 61.3% of TOEFL essays by non-native English writers as AI-generated, versus 3.2% of essays by US-born writers.

≈56% chancerechecking…

A ~4% per-sentence false-positive rate compounds: a 20-sentence paper has better-than-even odds of containing at least one falsely flagged sentence.

08

The Sorting Machine

71 vs 70 cutoffcited

A one-point difference at a diagnostic cutoff (a 71 against a threshold of 70) flips a child from eligible to ineligible, and the classification follows them through the record.

09

Accountability Tracker

25 × 11, 1.33M students✓ recomputed in CI

The audit covers 25 U.S. universities across an 11-dimension framework, roughly 1.33 million students in scope, and every dimension score is backed by a quoted policy excerpt with a source link.

12

The Warehouse

49 held optionscited

The federal state can be read as a portfolio of 49 unexercised options, most of them illegible to any evaluator that scores only observable output.

14

The Sorting Machine, Wartime Edition

~40% completecited

Ukraine's State Register of Property Rights is about 40% complete and mostly post-2013, so a pre-2013 paper-deed owner cannot prove ownership — the gate assisted filing cannot open.

19

Counting the Killings

84 to 146 more a yearrechecking…

The police-killing databases disagree by definition, not by error: Mapping Police Violence starts from the Washington Post's fatal shootings and adds the killings the Post's shooting-only rule excludes, so in every year from 2015 to 2024 MPV's count exceeds the Post's, by 84 to 146 people, the non-shooting police killings.

80% cleared only in 2024cited

The FBI's National Use-of-Force collection needs participating agencies to employ at least 80 percent of the country's sworn officers before it can publish a national total, a bar it first cleared in 2024, in its sixth year, then slipped back below.

>98% every yearrechecking…

The civilian databases converged: computed from Mapping Police Violence's own WaPo-ID and Fatal-Encounters-ID linkage, in every year the three lists overlap (2015 to 2021) more than 98 percent of MPV's killings also appear in the Post or Fatal Encounters. The disagreement is a thin, definitional residue, not disorder.

55.5% missed by the statecited

A two-source capture-recapture is invalid on these nested lists. The valid undercount estimate comes from an independent official source: the GBD study found US vital statistics failed to attribute 55.5 percent of police-violence deaths, 17,100 of an estimated 30,800 over 1980 to 2018, to the police.

+18% on both countersrechecking…

The count is rising under both living rules: aggregated from each project's primary dataset, the Washington Post's fatal-shootings count and Mapping Police Violence's all-killings count are each up 18 percent from their first full year to 2024, and 2024 is the highest year on both.

highest 59.4%, rows reconcilerechecking…

The certificate undercount is unevenly distributed, a measurement finding about whose deaths the official records failed to attribute: the GBD study's group figures show misclassification highest for Black Americans (59.4 percent) and pervasive for every group, and the four group rows sum back to the study's overall 17,100-of-30,800 within published rounding.

2.5x / 2.9x / 3.5x, three rulesrechecking…

The racial disparity is definition-proof: every counter in this series that publishes rates by race finds Black Americans killed at two to three and a half times the white rate, under three different rules and eras. The Post's shootings-only rates (6.1 vs 2.4 per million per year, 2015-2024), MPV's all-killings likelihood ratio (2.9x, 2024), and the Lancet/GBD reconstruction (0.69 vs 0.20 per 100,000 age-standardized, 1980-2019). Where an instrument publishes both rates, its ratio re-derives live.

penalty applied 0 times since 2013cited

A reporting mandate already exists and has never been enforced: the Death in Custody Reporting Act of 2013 requires states to report deaths in custody and authorizes a 10 percent cut to a noncompliant state's federal justice grants, a penalty the Department of Justice has never once applied; collection began four years after the statute's own deadline, per the DOJ Inspector General's 2018 review.

49% at best, then ~1,900 foundcited

The government tested the civilian method and it worked: the BJS-sponsored assessment found the Arrest-Related Deaths program captured, at best, 49 percent of law-enforcement homicides (collection suspended 2014), and the redesigned open-source media review then identified 1,348 potential arrest-related deaths in ten months, an estimated 1,900 for the year.

1 in 1,000 lifetime, peak ages 20-35cited

The disparity as a lifetime burden: about 1 in 1,000 Black men can expect to be killed by police over the life course, roughly 2.5 times the risk carried by white men, with risk peaking between ages 20 and 35, where police violence ranks among the leading causes of death for young Black men (Edwards, Lee & Esposito 2019, from Fatal Encounters data 2013-2018).

73 vs 458, one year, on definitioncited

The officer ledger is broken the same way: the FBI's felonious count (73 officers killed in 2021, the most since 2001) sits six times below NLEOMF's all-line-of-duty total for the same year (458, COVID the leading cause), the assault count is voluntary and covered about six in ten agencies in 2023, and the NIBRS transition collapsed reporting from about 16,000 agencies to roughly 5,000 in 2021.

70% incomplete, ~1,000 missed, one auditcited

The word 'mandated' is true at every layer separately and false about the system: department policy and state statutes like California's AB 71 (2015, all agencies, serious injury/discharge/death) and Texas's five-day officer-involved-shooting reports genuinely mandate reporting, while the GAO's 2022 audit of the one federal death mandate found 70 percent of state-submitted records missing a required element and nearly 1,000 deaths in one fiscal year potentially unreported, with DOJ yet to determine any state's compliance.

Computed claims are recomputed from the shipped models (src/lib/research) and the audit dataset; sourced figures are cross-checked against the essays in CI (src/lib/research/__tests__/claims.test.js). This page is the human-readable face of that guard.

Need something like this built?

I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.

Work with me →