Span tree and waterfall
What called what, in what order, and how long did each step take?
MonitorDebugLivePost-mortem
Nested spans on a shared time axis. The default in every observability product, now standardized by the OTel GenAI conventions as invoke_agent, chat and execute_tool.
- Strength
- Exact timing, cost and token attribution; portable across vendors.
- Limit
- Grows unreadable at hundreds of decisions; shows whether a step ran, not whether it was right.
- Sources
- OTel GenAI conventions, Honeycomb, Morph tracing guide
- Demo
- See it on the shared run
Execution graph
What shape is this agent system, and which path did this run take?
DebugExplainPost-mortemDesign
Agents or steps as nodes, control flow as edges. Either inferred from trace nesting or read from the framework's declared graph.
- Strength
- Matches the developer's mental model of the pipeline; aggregated views keep busy agents readable.
- Limit
- Aggregation hides per-iteration differences; unrolled views drift back toward a tree. Few tools explain why an edge was taken.
- Sources
- Langfuse, LangSmith Studio, Langfuse discussion
- Demo
- See it on the shared run
Hierarchical temporal summary
How did many agents' behavior evolve, and what caused this event?
ExplainDebugPost-mortem
LLM-summarized behavior grouped into levels over time, with per-agent trajectories and causal tracing back through prior events.
- Strength
- Handles long, many-agent runs; supports cause-finding rather than log reading.
- Limit
- Summaries are themselves model output and can be wrong; built for simulations more than task pipelines.
- Sources
- AgentLens
- Demo
- See it on the shared run
Layered semantic summary
Did the plan fail, did a step get skipped, or did a tool call go wrong?
DebugPost-mortemLive
The run is restated at several levels of intent, from overall goal down to concrete executions. DiLLS borrows activity, action and operation from Activity Theory; MUSE keeps a similar summary updating live.
- Strength
- Lets developers spot planning failures and skipped steps without reading transcripts.
- Limit
- DiLLS is static and post-mortem and assumes a central orchestrator. Live versions are new and lightly evaluated.
- Sources
- DiLLS, MUSE
- Demo
- See it on the shared run
Branching message history
If I change this message and rerun from here, what happens?
DebugSteerLivePost-mortem
The conversation between agents as a timeline users can fork: edit a prior message, restore checkpointed agent state, and compare branches in an overview.
- Strength
- Turns debugging into experiment; strong evidence from a CHI user study that resets are central.
- Limit
- Tied to one framework (AutoGen); overview visual is secondary to the message UI.
- Sources
- AGDebugger, AgentGUI
- Demo
- See it on the shared run
Coordination plan view
Who does what, depending on whom, before anything runs?
DesignDesign
A structured representation of roles, task dependencies and result correspondence, generated from a goal and edited visually before execution, then linked to results after.
- Strength
- Moves errors upstream to the specification, where MAST finds the largest share of failures.
- Limit
- Evaluated with 12 people on toy tasks; gives little visibility during a run.
- Sources
- AgentCoord, MAST
- Demo
- See it on the shared run
Cross-run comparison
Why did the same task behave differently on different runs?
DebugEvaluatePost-mortem
Aligns repeated executions of one task so divergence points and behavioral variance become visible, instead of inspecting one trace at a time.
- Strength
- Addresses non-determinism head on, which single-trace tools cannot.
- Limit
- Very little work so far; InconLens is still a preprint.
- Sources
- InconLens
- Demo
- See it on the shared run
Failure-mode overlay
Which known kind of failure is this, and where did it start?
EvaluateDebugPost-mortem
Annotating traces with a failure taxonomy, often via an LLM judge, so failure types can be counted, located and compared across systems.
- Strength
- Gives shared vocabulary and datasets; MAST's annotated traces are public.
- Limit
- Mostly produces tables and counts. Encoding failure modes directly in an interactive view is largely unbuilt.
- Sources
- MAST, Trajectory Analysis Survey
- Demo
- See it on the shared run