Loops get flattened
Agent pipelines are loops (plan, act, observe, retry), yet most views unroll them into trees or DAGs. Only Langfuse's aggregated mode draws cycles, and it discards per-iteration detail.
Evidence: Langfuse
Research · Interactive essay · No. 28
Four agents wrote a one-page brief. Every step reported success. The brief carries a claim nobody sourced. Here is that run drawn eight ways, each showing a different slice of what went wrong, and two that cannot see it at all.
When a multi-agent system fails, the first question is not what went wrong but where to look. The trace exists. It has a few dozen spans, each with a start, an end and a status, and every status says ok. Somewhere in that record an orchestrator decided to skip a step, a researcher wrote down a claim without a source, and a critic waved it through. Nothing in the record is red.
This post takes one run and draws it eight ways. The techniques come from a review of sixteen sources: papers, preprints and vendor documentation, all listed on the companion page. Each section says what question the technique answers, shows it on the same run, and says what it misses. The comparison is the argument: the same forty-odd spans, and what each picture lets you notice.
The run is synthetic, and it was built to fail in one specific way. Four agents share one goal, and every span they produce reports status ok. Read the cast first; every demo below is drawn from the same record.
Goal: Write a one-page brief on the history of the metric system
plantedSkips one planned step, the source check, because no agent owns it.
plantedOne claim, C3, has no citation.
plantedRevises once after a rejection, which is the loop.
plantedRejects v1 over an unrelated claim, then accepts v2, which still carries C3.
Every span reports status ok. That is the point. The spans are shaped like OpenTelemetry GenAI agent spans (invoke_agent, chat, execute_tool), and the run is labeled synthetic on every demo.
01 of 08
What called what, in what order, and how long did each step take?
This is the default. Every observability product draws it, and the OpenTelemetry GenAI conventions now say what the spans are called: an invoke_agent span for each agent turn, chat spans for the model calls inside it, execute_tool spans for the tool calls. Nesting is containment, the horizontal axis is time, and the color of a bar is its status.
On this run it is exact and unhelpful in the same moment. You can read that the researcher took seventeen seconds and that the critic checked four claims twice. You cannot read that one of those claims had no citation, because the check span that looked at it returned ok, and ok is all the bar knows.
What it misses. Correctness. The step that let the bad claim through is the same green as every other step. At four hundred spans the tree stops being readable at all.
Sources: OTel GenAI conventions, Honeycomb, Morph tracing guide.
02 of 08
What shape is this agent system, and which path did this run take?
Collapse time and keep control flow. Agents and tools become nodes; who handed off to whom becomes an edge. Langfuse infers this from span timing and nesting; LangSmith reads it from the declared graph. Aggregated mode counts repeats and draws the writer and critic as a cycle. Expanded mode unrolls every call into execution order.
The graph matches how the developer imagined the pipeline, and that is both its value and its trap. The aggregated view shows a loop, which the waterfall cannot, and tapping an edge shows the route reason recorded for it. It does not show that the loop went around for the wrong reason.
What it misses. Per-iteration difference. Aggregation says the critic ran twice; it does not say what changed between the two runs, and few tools record why an edge was taken at all.
Sources: Langfuse, LangSmith Studio, Langfuse discussion.
03 of 08
How did many agents' behavior evolve, and what caused this event?
AgentLens takes the opposite bet from the graph: keep time, add meaning. One swimlane per agent, events along it, and behind each event a chain of causes back through earlier events. Zoom out and the events are summarized; zoom in and they are the raw calls.
The causal chain is what earns the technique its place. Tap the critic's accept and the trace lights up backward: the second draft, the first rejection, the claims message, and the orchestrator's decision to move on without a source check. The uncited claim is a few hops from the accept.
What it misses. AgentLens summarizes with a model, and a model can summarize wrong. It was also built for agent simulations, where hundreds of agents drift, not for a four-agent task pipeline.
Sources: AgentLens.
04 of 08
Did the plan fail, did a step get skipped, or did a tool call go wrong?
DiLLS restates the run at three levels borrowed from Activity Theory: the activity (what the system was trying to do), the actions (the steps the plan committed to), and the operations (what actually executed). MUSE keeps a similar summary updating while the run is still going.
The skipped step lives at the action level. The plan said check sources; no action row ever landed for it, and the summary says so. This is the one technique on the page that names the failure in words, which is why it is also the one most dependent on its summaries being right.
What it misses. Anything the summarizer did not think to say. DiLLS is post-mortem and assumes an orchestrator; live versions are new and lightly evaluated.
05 of 08
If I change this message and rerun from here, what happens?
AGDebugger treats the conversation between agents as the primary object. Every message is a point you can return to: edit one, restore the agents' state from a checkpoint, and the run forks. A small overview keeps the branches in view.
Here the edit is the researcher's claims message. Add one line flagging the missing citation, rerun from that point, and the critic catches C3 on the second lap, the writer removes it on a third, and the accept span is finally honest. Only that one edit is live here; the branch is pre-computed, and the page says so.
What it misses. The overview is the secondary view; the messages are the primary one. AGDebugger is bound to one framework (AutoGen), so the technique travels less easily than the idea.
Sources: AGDebugger, AgentGUI.
06 of 08
Who does what, depending on whom, before anything runs?
AgentCoord moves the picture upstream of execution: a matrix of agents by tasks, a dependency graph between the tasks, and the ability to edit both before pressing go. MAST's finding that system design failures are the largest category is the argument for looking here first.
The plan for this run has a hole you can see without running anything. Check sources has no owner. Verify depends on draft but not on check, so even an owner would not have helped the critic. Reassign the task and the validator re-checks the dependencies in place.
What it misses. The run itself. A plan view shows what should happen; it says nothing about what did, and AgentCoord's evaluation was twelve people on toy tasks.
Sources: AgentCoord, MAST.
07 of 08
Why did the same task behave differently on different runs?
Run the same goal six times and the researcher's second search comes back with a different page twice. In those two runs the page happens to support C3, the claim gets a citation, and the failure never happens. InconLens is the first tool to make that the object of study: align repeated runs, mark where they diverge, show the variance step by step.
The point is that the failure is intermittent. A single-trace tool sees either a clean run or a bad one and cannot tell you the difference was a search result. Six aligned runs can.
What it misses. Almost everything, so far. InconLens is a preprint and there is little else; the technique is mostly a gap.
Sources: InconLens.
08 of 08
Which known kind of failure is this, and where did it start?
MAST annotated more than 1,600 traces and found fourteen ways multi-agent systems fail, in three categories: system design issues, inter-agent misalignment, and task verification. The taxonomy is usually a table. Here it is drawn onto the waterfall from the first section as marks.
Two of the fourteen are present. The accept span carries incorrect verification: the critic checked each claim for consistency with the list, not for a source. The plan span carries a system design failure: a step nobody owned. The marks are linked by the same causes the temporal view followed. Green bars, two annotations, and now the trace says what it means.
What it misses. The marks are labels, and someone or something had to apply them. MAST's annotations came from a model judge checked against people; encoding them directly in an interactive view is largely unbuilt.
Sources: MAST, Trajectory Analysis Survey.
Thesis
What changed between the last time around and this time?
Every view above unrolls the loop or collapses it. The waterfall lays the two laps end to end. The graph folds them into one cycle with a count of two. Neither answers the question a person debugging a retry actually asks: what changed between the last time around and this time?
So draw the loop as the primary shape. The writer and critic cycle is a closed track; each lap is one trip around it; the spans of that lap sit on the track where they happened. Scrub laps and the track redraws. Put lap 1 and lap 2 side by side and the diff is the story: the C2 check changed, the verdict changed, and the C3 check is identical, citation none, both times.
That last row is the whole failure, and only the loop view puts it in one line.
Sources: Langfuse, Morph tracing guide, MAST.
Eight pictures of one run. The failure was visible in six of them, two of which name it in words, and invisible in the two most widely deployed. That ratio is the finding. The literature behind the techniques knows it, too: below are the gaps the review found most promising, leading with the two this run made concrete.
Agent pipelines are loops (plan, act, observe, retry), yet most views unroll them into trees or DAGs. Only Langfuse's aggregated mode draws cycles, and it discards per-iteration detail.
Evidence: Langfuse
Traces report whether a span succeeded, not whether the step was right. MAST supplies a failure vocabulary, but tools rarely draw it.
Evidence: Morph tracing guide, MAST
Practitioner guides say span trees stop working around a few hundred decisions. Langfuse's aggregated mode is one fix; AgentLens uses LLM summaries. Neither has been compared against the other.
Evidence: Morph tracing guide, Langfuse, AgentLens
InconLens argues nearly all diagnosis tools examine one trace at a time, while agents behave differently on repeated runs.
Evidence: InconLens
AgentGUI finds that trajectory visualizers leave out steering and steering tools are bound to one harness. MUSE contrasts its live view with DiLLS's post-mortem one.
Evidence: AgentGUI, MUSE, DiLLS, OTel GenAI conventions
Nearly every system targets agent developers. Langfuse users explicitly asked for a view a non-technical support person could follow.
Evidence: Langfuse discussion, MUSE
Core systems report user studies of 12 to 14 participants, often on toy tasks. No shared benchmark for agent visualization exists.
Evidence: AgentCoord, AGDebugger, Trajectory Analysis Survey
The run is synthetic and labeled as such. Every summary, branch and failure label is pre-computed in a committed fixture and re-verified by tests on every change; nothing on this page calls a model. Claims about papers and tools map to the source ledger on the review page, where preprints are marked as not peer reviewed.
Full citations, the technique taxonomy with filters, and the review's own limits are on the companion review page. The run itself is downloadable as dataset.json (CC BY 4.0).
I design and ship AI tools, full-stack apps, and data pipelines — end to end, to production. Tell me the problem in a sentence; I'll give you an honest read on fit within a day.
Work with me →