Criterion and numbers
- Registered criterion
- SAGEN @ 300 tokens with live perception beats the strongest flat baseline on coverage
- Measured
- +0.230 coverage over the rolling summary
- 95% CI
- 0.205 to 0.257
- p
- 0.0001 (Holm threshold 0.0125)
SAGEN-BENCH V2 · PRE-REGISTERED
SAGEN gives AI agents a structured working memory of a conversation: goals, topics, what refers to what. Four predictions were filed on OSF before testing it on 72 frozen scenarios. Two held. Two failed, and they're published here as failures.
Coverage: how much of what the agent should be tracking actually made it into its memory, from 0 to 1.
| Approach | Oracle | Live | Gap |
|---|---|---|---|
| Raw buffer | 0.160 | 0.140 | 0.021 |
| Full transcript | 0.195 | 0.174 | 0.021 |
| Rolling summary | 0.268 | 0.268 | -0.000 |
| SAGEN @ 300 tokens | 0.937 | 0.499 | 0.438 |
| SAGEN @ 2000 tokens | 0.966 | 0.828 | 0.138 |
H3With perfect inputs SAGEN captures 0.937; when a real model does the reading, 0.499. The registered bar allowed a gap of 0.10. The gap was 0.438.
H4Models matched ground truth on about a quarter of what they perceived (alignment 0.257 against a 0.80 bar) and weren't consistent run to run (0.719 against 0.90).
Share of what each model perceived that matched the frozen ground truth. 5 runs over 30 scenarios per model.
| Model | Alignment | Trials | Role |
|---|---|---|---|
| claude-haiku-4-5-20251001 | 0.251 | 150 | sweep |
| claude-sonnet-4-6 | 0.257 | 150 | flagship, the H4 gate |
| claude-opus-4-8 | 0.273 | 150 | sweep |
A bigger model barely moves it.
The architecture is not the bottleneck; perception is. That's where the next work goes.
Each criterion below was written down and filed on OSF before the paid window ran. The confirmatory family is corrected with Holm-Bonferroni at alpha 0.05. Open a verdict to see its registered criterion and the measured value.
The Holm table marks H4 as significant because its two-sided p only detects distance from the floor. The mean sits below the floor, so the directional criterion fails. See the disclosures.
None of the four hypothesis verdicts use the judge; they are computed mechanically, so this gate failing does not touch them.
Before the confirmatory run, 64 predictions were filed: remove one SAGEN mechanism, and exactly these capability probes should break. Rows are the mechanism removed; columns are the capability probed.
This grid is the committed result on the frozen 72-scenario corpus, the one CI re-verifies on every commit.
PERSISTENCE ABLATION · ORACLE MODE
The knockout lattice removes mechanisms but never removes persistence itself. This free, deterministic ablation does. It compares the persistent engine against a stateless structured baseline: a fresh engine each turn that sees only that turn's analysis and injects it, carrying no memory forward. The per-dimension delta separates the coverage that persistence buys from the coverage the typed schema emits for free.
Coverage with the persistent engine minus coverage with a memoryless structured baseline, per dimension, at the 300-token budget. Right of zero: memory helps. Left of zero: it costs.
| Dimension | Persistence delta | Reads as |
|---|---|---|
| Explicit goals explicit-goal-identification | +1.000 | needs memory |
| Inferred goals inferred-goals | +0.972 | needs memory |
| Topic pivots topic-pivot-detection | +0.893 | needs memory |
| Active topics active-topic-tracking | +0.743 | needs memory |
| Goal priority goal-priority | +0.681 | needs memory |
| Goal lifecycle goal-lifecycle | +0.681 | needs memory |
| Memory decay memory-decay-compression | +0.587 | needs memory |
| Callbacks callback-detection | 0.000 | free from the schema |
| Sentiment sentiment-tracking | 0.000 | free from the schema |
| Sentiment urgency sentiment-urgency | 0.000 | free from the schema |
| Entity types entity-type-classification | 0.000 | free from the schema |
| Watchlist patterns scan-pattern-watchlist | 0.000 | free from the schema |
| Fits the token budget token-budget-rendering | 0.000 | free from the schema |
| Transition order temporal-transition-ordering | -0.028 | free from the schema |
| Transition types transition-type-classification | -0.056 | budget truncation artifact |
| Machine-parseable output machine-parseable-output | -0.215 | budget truncation artifact |
Persistence is load-bearing. It carries goal tracking, topic-pivot detection, cumulative topics and memory decay. But as many dimensions come free with no memory at all: callback and sentiment ride in the per-turn analysis, and fitting the token budget is a format property. So part of SAGEN's advantage over flat memory is schema affordance, not knowing more. The two dimensions left of zero are a truncation artifact of carrying more state through the same 300 tokens. That is the honest reading of a coverage number. This is oracle mode; the live-perception ablation is future work.
DISCLOSURES · DEVIATIONS AND KNOWN GAPS
Everything the window did differently from the registration, and every gap in the instruments. Open any one.
Registered: S2 temperature-0 and sampled cells kept separate; run-to-run stability measured at temperature 0.
Actual: Current Claude API tiers reject an explicit temperature parameter, so all runs execute at each model's default sampling and the R = 5 runs form ONE cell per model.
Bias: Makes H4's run-to-run stability floor STRICTER (default sampling includes sampling noise that temperature 0 would have suppressed); H4 can fail from this deviation but cannot spuriously pass because of it.
Registered: Model ids frozen at registration.
Actual: The registration text names roles and counts but not vendor model ids; ids were pinned in this committed file before the first paid call.
Bias: None on outcomes; a completeness gap in the registration transcription, disclosed.
The Holm table marks H4 "reject H0" from a two-sided bootstrap p that only detects distance from the 0.80 floor. The alignment mean sits far below the floor, so the directional registered criterion fails: the verdict for H4 is FAILS, as shown above.
935 of 960 realized turns passed the frozen fidelity checks outright. 25 fell back to logged, frozen template text after 3 failed rewrite attempts, and on 7 of those the template itself trips the spurious-callback cue. An instrument gap in the cue list, not a data problem: the affected turns carry ground-truth text verbatim.
Every number above re-derives from a public bundle on your own machine: the frozen corpus, the raw record of every paid call, and the instruments that scored them. No account, no API key, no dependencies.
Download the bundle and run the verifier (Node 20 or newer)
B=https://www.jakelawrence.xyz/downloads/sagen-bench/v2; mkdir sagen-bench-v2 && cd sagen-bench-v2 && curl -fsSO $B/SHA256SUMS && awk '{print $2}' SHA256SUMS | xargs -I{} curl -fsS --create-dirs -o {} $B/{} && node verify.mjsIt fetches every file listed in SHA256SUMS, re-derives the corpus hash, the lattice, the coverage means, H1 to H4 with their bootstrap intervals, the Holm family, the judge gate and the spend, prints PASS for each, and ends on VERIFIED.
I build the AI tools and automations a team adopts, and the full-stack apps and data pipelines under them, and ship them to production. Tell me the problem in a sentence and I'll give you an honest read on fit within a day.
Work with me →