Browse documentation

Investigate

Investigate agent runs

Triage AI executions by customer, status, cost, tools, and quality evidence.

Open Explore → Agent Runs to investigate an AI workflow as one execution rather than a collection of unrelated model spans.

Find the run that matters

Filter by agent, customer or distinct ID, run status, and quality state. Sort by duration, steps, tool calls, or cost to find outliers. A run without a customer link usually means identity was not active before its first span.

Open a run and review:

  • the execution sequence and slow or failed steps;
  • repeated model or tool calls;
  • token use and calculated cost;
  • errors and status transitions;
  • human scores and automatic evaluation results; and
  • connected customer, trace, replay, or Issue evidence.

Interpret quality scores carefully

Human scores capture reviewer judgment. Automatic scores apply the organization’s configured sample and criteria. Neither score explains a failure by itself; compare the cited run steps, tool outputs, and application evidence.

If a run was not sampled, that is not a passing score. If an evaluation lacks the inputs needed for a criterion, change the recorded evidence or criterion before increasing the sample rate.

Move from run to impact

Open the customer to see what preceded and followed the run. Compare similar runs for other customers before treating one expensive or failed execution as systemic. Use Customer Detective when a grounded cross-signal explanation is more useful than manually reading every span.