Research · Interactive essay · No. 28

Eight Ways to See an Agent Think

Four agents wrote a one-page brief. Every step reported success. The brief carries a claim nobody sourced. Here is that run drawn eight ways, each showing a different slice of what went wrong, and two that cannot see it at all.

The shared run, 40 spans in 52 seconds, colored by agent. Every status is ok. The marks are the two failures no status carries.
8 techniques40 spans in the run6 repeat runs16 sources0 model calls on this page
Source ledger review pagedataset.json the run

When a multi-agent system fails, the first question is not what went wrong but where to look. The trace exists. It has a few dozen spans, each with a start, an end and a status, and every status says ok. Somewhere in that record an orchestrator decided to skip a step, a researcher wrote down a claim without a source, and a critic waved it through. Nothing in the record is red.

This post takes one run and draws it eight ways. The techniques come from a review of sixteen sources: papers, preprints and vendor documentation, all listed on the companion page. Each section says what question the technique answers, shows it on the same run, and says what it misses. The comparison is the argument: the same forty-odd spans, and what each picture lets you notice.

The shared run

The run is synthetic, and it was built to fail in one specific way. Four agents share one goal, and every span they produce reports status ok. Read the cast first; every demo below is drawn from the same record.

Goal: Write a one-page brief on the history of the metric system

OrchestratorPlans three subtasks from the goal, assigns them, collects results.

plantedSkips one planned step, the source check, because no agent owns it.

ResearcherRuns two searches and two fetches, returns four claims with citations.

plantedOne claim, C3, has no citation.

WriterDrafts the brief from the claims.

plantedRevises once after a rejection, which is the loop.

CriticChecks the draft against the claims list.

plantedRejects v1 over an unrelated claim, then accepts v2, which still carries C3.

Every span reports status ok. That is the point. The spans are shaped like OpenTelemetry GenAI agent spans (invoke_agent, chat, execute_tool), and the run is labeled synthetic on every demo.

01 of 08

Span waterfall

What called what, in what order, and how long did each step take?

This is the default. Every observability product draws it, and the OpenTelemetry GenAI conventions now say what the spans are called: an invoke_agent span for each agent turn, chat spans for the model calls inside it, execute_tool spans for the tool calls. Nesting is containment, the horizontal axis is time, and the color of a bar is its status.

On this run it is exact and unhelpful in the same moment. You can read that the researcher took seventeen seconds and that the critic checked four claims twice. You cannot read that one of those claims had no citation, because the check span that looked at it returned ok, and ok is all the bar knows.

Demo 1: Span tree and waterfallsynthetic run
Loading the demo
Try thisExpand the critic's second turn and read the C3 check. Then toggle the scale test.

What it misses. Correctness. The step that let the bad claim through is the same green as every other step. At four hundred spans the tree stops being readable at all.

Sources: OTel GenAI conventions, Honeycomb, Morph tracing guide.

02 of 08

Execution graph

What shape is this agent system, and which path did this run take?

Collapse time and keep control flow. Agents and tools become nodes; who handed off to whom becomes an edge. Langfuse infers this from span timing and nesting; LangSmith reads it from the declared graph. Aggregated mode counts repeats and draws the writer and critic as a cycle. Expanded mode unrolls every call into execution order.

The graph matches how the developer imagined the pipeline, and that is both its value and its trap. The aggregated view shows a loop, which the waterfall cannot, and tapping an edge shows the route reason recorded for it. It does not show that the loop went around for the wrong reason.

Demo 2: Execution graphsynthetic run
Loading the demo
Try thisToggle Aggregated and Expanded. Tap the edge from critic to writer.

What it misses. Per-iteration difference. Aggregation says the critic ran twice; it does not say what changed between the two runs, and few tools record why an edge was taken at all.

Sources: Langfuse, LangSmith Studio, Langfuse discussion.

03 of 08

Hierarchical temporal summary

How did many agents' behavior evolve, and what caused this event?

AgentLens takes the opposite bet from the graph: keep time, add meaning. One swimlane per agent, events along it, and behind each event a chain of causes back through earlier events. Zoom out and the events are summarized; zoom in and they are the raw calls.

The causal chain is what earns the technique its place. Tap the critic's accept and the trace lights up backward: the second draft, the first rejection, the claims message, and the orchestrator's decision to move on without a source check. The uncited claim is a few hops from the accept.

Demo 3: Hierarchical temporal summarysynthetic run
Loading the demo
Try thisTap the accept event at the end of the critic's lane and follow the chain.

What it misses. AgentLens summarizes with a model, and a model can summarize wrong. It was also built for agent simulations, where hundreds of agents drift, not for a four-agent task pipeline.

Sources: AgentLens.

04 of 08

Layered semantic summary

Did the plan fail, did a step get skipped, or did a tool call go wrong?

DiLLS restates the run at three levels borrowed from Activity Theory: the activity (what the system was trying to do), the actions (the steps the plan committed to), and the operations (what actually executed). MUSE keeps a similar summary updating while the run is still going.

The skipped step lives at the action level. The plan said check sources; no action row ever landed for it, and the summary says so. This is the one technique on the page that names the failure in words, which is why it is also the one most dependent on its summaries being right.

Demo 4: Layered semantic summarysynthetic run
Loading the demo
Try thisPress Play and watch the action rows land. The flagged row is the step that never ran.

What it misses. Anything the summarizer did not think to say. DiLLS is post-mortem and assumes an orchestrator; live versions are new and lightly evaluated.

Sources: DiLLS, MUSE.

05 of 08

Branching message history

If I change this message and rerun from here, what happens?

AGDebugger treats the conversation between agents as the primary object. Every message is a point you can return to: edit one, restore the agents' state from a checkpoint, and the run forks. A small overview keeps the branches in view.

Here the edit is the researcher's claims message. Add one line flagging the missing citation, rerun from that point, and the critic catches C3 on the second lap, the writer removes it on a third, and the accept span is finally honest. Only that one edit is live here; the branch is pre-computed, and the page says so.

Demo 5: Branching message historysynthetic run
Loading the demo
Try thisEdit the researcher's message and rerun. Compare the two branches' lap counts.

What it misses. The overview is the secondary view; the messages are the primary one. AGDebugger is bound to one framework (AutoGen), so the technique travels less easily than the idea.

Sources: AGDebugger, AgentGUI.

06 of 08

Coordination plan view

Who does what, depending on whom, before anything runs?

AgentCoord moves the picture upstream of execution: a matrix of agents by tasks, a dependency graph between the tasks, and the ability to edit both before pressing go. MAST's finding that system design failures are the largest category is the argument for looking here first.

The plan for this run has a hole you can see without running anything. Check sources has no owner. Verify depends on draft but not on check, so even an owner would not have helped the critic. Reassign the task and the validator re-checks the dependencies in place.

Demo 6: Coordination plan viewsynthetic run
Loading the demo
Try thisGive Check sources to the critic. Then make Verify depend on it.

What it misses. The run itself. A plan view shows what should happen; it says nothing about what did, and AgentCoord's evaluation was twelve people on toy tasks.

Sources: AgentCoord, MAST.

07 of 08

Cross-run comparison

Why did the same task behave differently on different runs?

Run the same goal six times and the researcher's second search comes back with a different page twice. In those two runs the page happens to support C3, the claim gets a citation, and the failure never happens. InconLens is the first tool to make that the object of study: align repeated runs, mark where they diverge, show the variance step by step.

The point is that the failure is intermittent. A single-trace tool sees either a clean run or a bad one and cannot tell you the difference was a search result. Six aligned runs can.

Demo 7: Cross-run comparisonsynthetic run
Loading the demo
Try thisTap the divergence marker at the second search, then the claims step.

What it misses. Almost everything, so far. InconLens is a preprint and there is little else; the technique is mostly a gap.

Sources: InconLens.

08 of 08

Failure-mode overlay

Which known kind of failure is this, and where did it start?

MAST annotated more than 1,600 traces and found fourteen ways multi-agent systems fail, in three categories: system design issues, inter-agent misalignment, and task verification. The taxonomy is usually a table. Here it is drawn onto the waterfall from the first section as marks.

Two of the fourteen are present. The accept span carries incorrect verification: the critic checked each claim for consistency with the list, not for a source. The plan span carries a system design failure: a step nobody owned. The marks are linked by the same causes the temporal view followed. Green bars, two annotations, and now the trace says what it means.

Demo 8: Failure-mode overlaysynthetic run
Loading the demo
Try thisFilter by category. Count how many modes are present against how many exist.

What it misses. The marks are labels, and someone or something had to apply them. MAST's annotations came from a model judge checked against people; encoding them directly in an interactive view is largely unbuilt.

Sources: MAST, Trajectory Analysis Survey.

Thesis

The loop is the unit

What changed between the last time around and this time?

Every view above unrolls the loop or collapses it. The waterfall lays the two laps end to end. The graph folds them into one cycle with a count of two. Neither answers the question a person debugging a retry actually asks: what changed between the last time around and this time?

So draw the loop as the primary shape. The writer and critic cycle is a closed track; each lap is one trip around it; the spans of that lap sit on the track where they happened. Scrub laps and the track redraws. Put lap 1 and lap 2 side by side and the diff is the story: the C2 check changed, the verdict changed, and the C3 check is identical, citation none, both times.

That last row is the whole failure, and only the loop view puts it in one line.

Thesis demo: loop viewsynthetic run
Loading the demo
Try thisScrub to lap 2, then compare laps 1 and 2.

Sources: Langfuse, Morph tracing guide, MAST.

What is still open

Eight pictures of one run. The failure was visible in six of them, two of which name it in words, and invisible in the two most widely deployed. That ratio is the finding. The literature behind the techniques knows it, too: below are the gaps the review found most promising, leading with the two this run made concrete.

Loops get flattened

Agent pipelines are loops (plan, act, observe, retry), yet most views unroll them into trees or DAGs. Only Langfuse's aggregated mode draws cycles, and it discards per-iteration detail.

Research angleLoop-native encodings where the cycle is the primary shape and iterations are layered on it. Game interfaces built around a visible, repeating loop are an untested design source here.

Evidence: Langfuse

Structure is shown, correctness is not

Traces report whether a span succeeded, not whether the step was right. MAST supplies a failure vocabulary, but tools rarely draw it.

Research angleA trace view that encodes MAST-style failure modes as first-class marks, validated against MAST-Data.

Evidence: Morph tracing guide, MAST

Scale breaks the dominant encoding

Practitioner guides say span trees stop working around a few hundred decisions. Langfuse's aggregated mode is one fix; AgentLens uses LLM summaries. Neither has been compared against the other.

Research angleControlled comparison of aggregation strategies (structural collapse vs. semantic summary) on long real-world traces.

Evidence: Morph tracing guide, Langfuse, AgentLens

Cross-run variance is nearly unstudied

InconLens argues nearly all diagnosis tools examine one trace at a time, while agents behave differently on repeated runs.

Research angleVisual alignment of N runs of the same pipeline to show where and why they diverge.

Evidence: InconLens

Observing and steering rarely meet

AgentGUI finds that trajectory visualizers leave out steering and steering tools are bound to one harness. MUSE contrasts its live view with DiLLS's post-mortem one.

Research angleA harness-neutral live view built on OTel GenAI spans that also accepts interventions.

Evidence: AgentGUI, MUSE, DiLLS, OTel GenAI conventions

Developers are the only audience

Nearly every system targets agent developers. Langfuse users explicitly asked for a view a non-technical support person could follow.

Research angleRepresentations for operators, managers or affected users, and what each audience needs to trust or contest an agent outcome.

Evidence: Langfuse discussion, MUSE

Evaluation is thin

Core systems report user studies of 12 to 14 participants, often on toy tasks. No shared benchmark for agent visualization exists.

Research angleA task battery for agent-trace comprehension, possibly seeded from MAST-Data or the TSE survey's benchmarks.

Evidence: AgentCoord, AGDebugger, Trajectory Analysis Survey

Method note

The run is synthetic and labeled as such. Every summary, branch and failure label is pre-computed in a committed fixture and re-verified by tests on every change; nothing on this page calls a model. Claims about papers and tools map to the source ledger on the review page, where preprints are marked as not peer reviewed.

Full citations, the technique taxonomy with filters, and the review's own limits are on the companion review page. The run itself is downloadable as dataset.json (CC BY 4.0).

Need something like this built?

I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.

Work with me →