SAGEN-BENCH V2 · PRE-REGISTERED

The memory works. Perception is the bottleneck.

SAGEN gives AI agents a structured working memory of a conversation: goals, topics, what refers to what. Four predictions were filed on OSF before testing it on 72 frozen scenarios. Two held. Two failed, and they're published here as failures.

Figure 1
What each memory approach keeps, with perfect inputs and with a real model reading

Coverage: how much of what the agent should be tracking actually made it into its memory, from 0 to 1.

Oracle mode: the agent is handed a perfect reading of the conversation. The ceiling. Live mode: a real model (Sonnet 4.6) reads the conversation itself. The realistic case. 72 scenarios in each mode; the frozen power analysis called for 66.
Figure 1 as a table
ApproachOracleLiveGap
Raw buffer0.1600.1400.021
Full transcript0.1950.1740.021
Rolling summary0.2680.268-0.000
SAGEN @ 300 tokens0.9370.4990.438
SAGEN @ 2000 tokens0.9660.8280.138
OSF DOI 10.17605/OSF.IO/S3GFM ↗72 scenarios8,488 model calls$84.75 of a $150 ceiling

WHAT FAILED, AND WHY

Two predictions failed. Both point at the same place.

H3With perfect inputs SAGEN captures 0.937; when a real model does the reading, 0.499. The registered bar allowed a gap of 0.10. The gap was 0.438.

H4Models matched ground truth on about a quarter of what they perceived (alignment 0.257 against a 0.80 bar) and weren't consistent run to run (0.719 against 0.90).

Figure 2
Perception alignment by model

Share of what each model perceived that matched the frozen ground truth. 5 runs over 30 scenarios per model.

Figure 2 as a table
ModelAlignmentTrialsRole
claude-haiku-4-5-202510010.251150sweep
claude-sonnet-4-60.257150flagship, the H4 gate
claude-opus-4-80.273150sweep

A bigger model barely moves it.

The architecture is not the bottleneck; perception is. That's where the next work goes.

REGISTERED VERDICTS

Every verdict, as filed

Each criterion below was written down and filed on OSF before the paid window ran. The confirmatory family is corrected with Holm-Bonferroni at alpha 0.05. Open a verdict to see its registered criterion and the measured value.

Registered hypotheses

H1Does SAGEN beat the best simple memory when a real model does the reading?HOLDS
Criterion and numbers
Registered criterion
SAGEN @ 300 tokens with live perception beats the strongest flat baseline on coverage
Measured
+0.230 coverage over the rolling summary
95% CI
0.205 to 0.257
p
0.0001 (Holm threshold 0.0125)
H2Is there structure that even a full transcript cannot hold?HOLDS
Criterion and numbers
Registered criterion
7 structural dimensions stay uncaptured by a full transcript in every scenario
Measured
0 violations across 72 scenarios
p
0.0001 (Holm threshold 0.0167)
H3Does a real model doing the reading cost little coverage?FAILS
Criterion and numbers
Registered criterion
Live coverage stays within 0.10 of oracle coverage
Measured
perception tax 0.438
95% CI
0.416 to 0.460
p
1.0000 (Holm threshold 0.0500)
H4Does the model read the conversation accurately and consistently?FAILS
Criterion and numbers
Registered criterion
Alignment at least 0.80 and run-to-run stability at least 0.90
Measured
alignment 0.257; stability 0.719
p
0.0001 (Holm threshold 0.0250)

The Holm table marks H4 as significant because its two-sided p only detects distance from the floor. The mean sits below the floor, so the directional criterion fails. See the disclosures.

Secondary

S3Does removing each mechanism break exactly what was predicted?HOLDS
Criterion and numbers
Registered criterion
All 64 knockout-by-probe cells behave as predicted in advance
Measured
64/64 on the frozen 72-scenario corpus

Quality gate

IRRDo two independent judge models agree?FAILS
Criterion and numbers
Registered criterion
Two judge models reach Cohen's kappa of at least 0.70
Measured
kappa 0.253 over 400 paired judgments

None of the four hypothesis verdicts use the judge; they are computed mechanically, so this gate failing does not touch them.

S3 · KNOCKOUT LATTICE

64 of 64 predictions came true

Before the confirmatory run, 64 predictions were filed: remove one SAGEN mechanism, and exactly these capability probes should break. Rows are the mechanism removed; columns are the capability probed.

  1. P1Pivot detection
  2. P2Progress continuity
  3. P3Callback opportunity
  4. P4Frustration threat
  5. P5Goal persistence
  6. P6Inferred goal capture
  7. P7Attention expiry
  8. P8Injection structure

This grid is the committed result on the frozen 72-scenario corpus, the one CI re-verifies on every commit.

PERSISTENCE ABLATION · ORACLE MODE

Does the blackboard earn its keep?

The knockout lattice removes mechanisms but never removes persistence itself. This free, deterministic ablation does. It compares the persistent engine against a stateless structured baseline: a fresh engine each turn that sees only that turn's analysis and injects it, carrying no memory forward. The per-dimension delta separates the coverage that persistence buys from the coverage the typed schema emits for free.

+0.338mean coverage from persistence (0.937 persistent vs 0.599 stateless, n = 72)
Figure 3
What persistence adds, dimension by dimension

Coverage with the persistent engine minus coverage with a memoryless structured baseline, per dimension, at the 300-token budget. Right of zero: memory helps. Left of zero: it costs.

Figure 3 as a table
DimensionPersistence deltaReads as
Explicit goals explicit-goal-identification+1.000needs memory
Inferred goals inferred-goals+0.972needs memory
Topic pivots topic-pivot-detection+0.893needs memory
Active topics active-topic-tracking+0.743needs memory
Goal priority goal-priority+0.681needs memory
Goal lifecycle goal-lifecycle+0.681needs memory
Memory decay memory-decay-compression+0.587needs memory
Callbacks callback-detection0.000free from the schema
Sentiment sentiment-tracking0.000free from the schema
Sentiment urgency sentiment-urgency0.000free from the schema
Entity types entity-type-classification0.000free from the schema
Watchlist patterns scan-pattern-watchlist0.000free from the schema
Fits the token budget token-budget-rendering0.000free from the schema
Transition order temporal-transition-ordering-0.028free from the schema
Transition types transition-type-classification-0.056budget truncation artifact
Machine-parseable output machine-parseable-output-0.215budget truncation artifact

Persistence is load-bearing. It carries goal tracking, topic-pivot detection, cumulative topics and memory decay. But as many dimensions come free with no memory at all: callback and sentiment ride in the per-turn analysis, and fitting the token budget is a format property. So part of SAGEN's advantage over flat memory is schema affordance, not knowing more. The two dimensions left of zero are a truncation artifact of carrying more state through the same 300 tokens. That is the honest reading of a coverage number. This is oracle mode; the live-perception ablation is future work.

DISCLOSURES · DEVIATIONS AND KNOWN GAPS

4 disclosures

Everything the window did differently from the registration, and every gap in the instruments. Open any one.

no-temperature-control

Registered: S2 temperature-0 and sampled cells kept separate; run-to-run stability measured at temperature 0.

Actual: Current Claude API tiers reject an explicit temperature parameter, so all runs execute at each model's default sampling and the R = 5 runs form ONE cell per model.

Bias: Makes H4's run-to-run stability floor STRICTER (default sampling includes sampling noise that temperature 0 would have suppressed); H4 can fail from this deviation but cannot spuriously pass because of it.

model-ids-pinned-at-window

Registered: Model ids frozen at registration.

Actual: The registration text names roles and counts but not vendor model ids; ids were pinned in this committed file before the first paid call.

Bias: None on outcomes; a completeness gap in the registration transcription, disclosed.

H4 Holm labeling

The Holm table marks H4 "reject H0" from a two-sided bootstrap p that only detects distance from the 0.80 floor. The alignment mean sits far below the floor, so the directional registered criterion fails: the verdict for H4 is FAILS, as shown above.

Realization fidelity

935 of 960 realized turns passed the frozen fidelity checks outright. 25 fell back to logged, frozen template text after 3 failed rewrite attempts, and on 7 of those the template itself trips the spurious-callback cue. An instrument gap in the cue list, not a data problem: the affected turns carry ground-truth text verbatim.

VERIFY IT YOURSELF

Don't take the page's word for it

Every number above re-derives from a public bundle on your own machine: the frozen corpus, the raw record of every paid call, and the instruments that scored them. No account, no API key, no dependencies.

Download the bundle and run the verifier (Node 20 or newer)

B=https://www.jakelawrence.xyz/downloads/sagen-bench/v2; mkdir sagen-bench-v2 && cd sagen-bench-v2 && curl -fsSO $B/SHA256SUMS && awk '{print $2}' SHA256SUMS | xargs -I{} curl -fsS --create-dirs -o {} $B/{} && node verify.mjs

It fetches every file listed in SHA256SUMS, re-derives the corpus hash, the lattice, the coverage means, H1 to H4 with their bootstrap intervals, the Holm family, the judge gate and the spend, prints PASS for each, and ends on VERIFIED.

GO DEEPER

Need something like this built?

I build the AI tools and automations a team adopts, and the full-stack apps and data pipelines under them, and ship them to production. Tell me the problem in a sentence and I'll give you an honest read on fit within a day.

Work with me →