Literature review · companion to the post

Visualizing LLM multi-agent pipelines

How researchers and tool builders make orchestrator, worker, tool-call and handoff behavior legible. Organized as a taxonomy of visual encodings, the sources behind them, and the open gaps worth a next study. Each encoding has a live demo on the post.

8 encodings16 sources7 peer-reviewed3 preprints6 industry7 gaps

Techniques

Purpose
When

8 of 8 encodings

Span tree and waterfall

What called what, in what order, and how long did each step take?

MonitorDebugLivePost-mortem

Nested spans on a shared time axis. The default in every observability product, now standardized by the OTel GenAI conventions as invoke_agent, chat and execute_tool.

Strength
Exact timing, cost and token attribution; portable across vendors.
Limit
Grows unreadable at hundreds of decisions; shows whether a step ran, not whether it was right.
Sources
OTel GenAI conventions, Honeycomb, Morph tracing guide
Demo
See it on the shared run

Execution graph

What shape is this agent system, and which path did this run take?

DebugExplainPost-mortemDesign

Agents or steps as nodes, control flow as edges. Either inferred from trace nesting or read from the framework's declared graph.

Strength
Matches the developer's mental model of the pipeline; aggregated views keep busy agents readable.
Limit
Aggregation hides per-iteration differences; unrolled views drift back toward a tree. Few tools explain why an edge was taken.
Sources
Langfuse, LangSmith Studio, Langfuse discussion
Demo
See it on the shared run

Hierarchical temporal summary

How did many agents' behavior evolve, and what caused this event?

ExplainDebugPost-mortem

LLM-summarized behavior grouped into levels over time, with per-agent trajectories and causal tracing back through prior events.

Strength
Handles long, many-agent runs; supports cause-finding rather than log reading.
Limit
Summaries are themselves model output and can be wrong; built for simulations more than task pipelines.
Sources
AgentLens
Demo
See it on the shared run

Layered semantic summary

Did the plan fail, did a step get skipped, or did a tool call go wrong?

DebugPost-mortemLive

The run is restated at several levels of intent, from overall goal down to concrete executions. DiLLS borrows activity, action and operation from Activity Theory; MUSE keeps a similar summary updating live.

Strength
Lets developers spot planning failures and skipped steps without reading transcripts.
Limit
DiLLS is static and post-mortem and assumes a central orchestrator. Live versions are new and lightly evaluated.
Sources
DiLLS, MUSE
Demo
See it on the shared run

Branching message history

If I change this message and rerun from here, what happens?

DebugSteerLivePost-mortem

The conversation between agents as a timeline users can fork: edit a prior message, restore checkpointed agent state, and compare branches in an overview.

Strength
Turns debugging into experiment; strong evidence from a CHI user study that resets are central.
Limit
Tied to one framework (AutoGen); overview visual is secondary to the message UI.
Sources
AGDebugger, AgentGUI
Demo
See it on the shared run

Coordination plan view

Who does what, depending on whom, before anything runs?

DesignDesign

A structured representation of roles, task dependencies and result correspondence, generated from a goal and edited visually before execution, then linked to results after.

Strength
Moves errors upstream to the specification, where MAST finds the largest share of failures.
Limit
Evaluated with 12 people on toy tasks; gives little visibility during a run.
Sources
AgentCoord, MAST
Demo
See it on the shared run

Cross-run comparison

Why did the same task behave differently on different runs?

DebugEvaluatePost-mortem

Aligns repeated executions of one task so divergence points and behavioral variance become visible, instead of inspecting one trace at a time.

Strength
Addresses non-determinism head on, which single-trace tools cannot.
Limit
Very little work so far; InconLens is still a preprint.
Sources
InconLens
Demo
See it on the shared run

Failure-mode overlay

Which known kind of failure is this, and where did it start?

EvaluateDebugPost-mortem

Annotating traces with a failure taxonomy, often via an LLM judge, so failure types can be counted, located and compared across systems.

Strength
Gives shared vocabulary and datasets; MAST's annotated traces are public.
Limit
Mostly produces tables and counts. Encoding failure modes directly in an interactive view is largely unbuilt.
Sources
MAST, Trajectory Analysis Survey
Demo
See it on the shared run

Need something like this built?

I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.

Work with me →