← Articles

Agent reliability · Practical guide

How to Detect AI Agent Failures Automatically from Traces

Most agent failures never throw an error. A practical procedure for finding them in production traces, and an honest look at what automated detection can and cannot do.

Faiz Vadakkumpadath
· 10 minute read
Smooth blue trajectories with a few breaking into orange loops and dead ends, each break circled

To detect AI agent failures automatically, stop looking for errors and start comparing runs. Group your production traces by the job the user was trying to get done, label each run as succeeded or failed using the signals you already have, and then look for what the failing runs do that the successful ones do not. Deterministic checks catch the obvious breakages, an LLM judge catches the semantic ones, and cohort comparison tells you which of those failures actually matter.

That is the short answer. The rest of this guide explains why each piece is necessary, what each detection method misses, and how to turn a pile of traces into a short list of failures worth fixing.

Why your error rate says everything is fine

Most agent failures do not throw exceptions. The model returns a well-formed answer, every tool call gets a 200, and the run finishes in a reasonable time. Your dashboards stay green.

Meanwhile the agent told a customer their booking was changed when it never called the booking API. Or it retried the same search six times and gave up. Or it answered a refund question with the baggage policy. None of these show up as errors, because from the system’s point of view nothing went wrong.

In our analysis of the public τ²-bench airline benchmark, 22% of runs told the customer the job was done without ever calling the API that would have done it. We wrote about that in how to evaluate AI agents in production. The final message looked perfect. Only the trace showed what actually happened.

An agent failure is usually a wrong decision, not a crash. You have to read the decisions to find it.

The four ways to detect agent failures, and what each one misses

There are four broad approaches. Teams usually start with the first and slowly discover they need the others.

MethodWhat it catchesWhat it misses
Deterministic checksTool errors, schema violations, timeouts, loops past a step limit, required calls that never happened, policy limits that were crossed.Anything that requires understanding language: a wrong answer, an unhelpful one, or a correct action taken for the wrong request.
Statistical anomaly detectionSudden changes in latency, token use, cost, step count, or tool mix after a deploy or a shift in traffic.Failures that have always been there. If 15% of refund runs have been quietly failing since launch, nothing about them looks anomalous.
LLM judge on each traceSemantic failures: hallucinated claims, ignored instructions, a skipped verification step, an answer that does not address the request.Patterns across runs. A judge sees one trace at a time and cannot tell you whether a failure is rare or happening to a third of your users.
Cohort comparisonThe decisions that separate failing runs from successful runs of the same job, and how many runs each one affects.Very rare one-off failures, and anything you have not labeled an outcome for.

None of these is enough on its own. Deterministic checks are cheap and precise but narrow. A judge is broad but noisy and expensive. Cohort comparison is what turns individual detections into a pattern you can prioritize. The rest of this guide combines them.

A step-by-step procedure that works on real traces

01

Split traces by the job the user wanted done

Production agents do not do one thing. An airline support agent handles bookings, cancellations, changes, baggage questions, insurance, and refunds. A failure in the refund flow has nothing to do with a failure in the booking flow, and the fix for one will not fix the other.

So the first step is classification. Read each trace and decide which business use case it was serving. If you skip this, every later statistic is an average over unrelated jobs, and real problems in small but important workflows disappear inside it. It is also the step Papaya automates first: every trace is classified into a business use case before any other analysis runs.

02

Label outcomes with the signals you already have

Automatic detection needs some notion of what success looks like. You rarely need hand labels to get started. Use the strongest evidence available for each workflow: the final database or ticket state, whether the user had to ask again, thumbs-down feedback, escalations to a human, a reopened ticket, or a support complaint tied back to the run.

Where none of those exist, a judge with a workflow-specific rubric can supply a provisional label. Treat it as provisional. More on that below.

03

Run cheap deterministic detectors on every trace

Before spending money on model calls, check the things code can check. Did a required tool get called before the agent claimed success? Did any tool return an error that the agent ignored? Did the same call repeat with identical arguments? Did a write happen without the confirmation step your policy requires? Did a tool return far more data than the next step could use?

Typical signals repeated identical tool calls, tool errors followed by a success message, missing required calls, schema mismatches, step-count and token outliers, writes without confirmation.

04

Judge the semantic failures, per workflow

Use an LLM judge only for questions code cannot answer, and give it a rubric written for that use case. “Was this a good response?” produces noise. “Did the agent confirm the fare difference before changing the booking?” produces a signal you can count. We cover how to build a judge you can trust in LLM-as-a-judge for AI agents.

05

Compare failing runs with successful runs of the same job

This is the step most teams skip, and it is the one that finds causes rather than symptoms. Take the failing and succeeding cohorts within one use case, then cluster inside each cohort. Look for the first point where they diverge: a different tool, a missing lookup, a much longer context, a sub-agent that returned twenty pages instead of two lines.

A difference that shows up in most failing runs and rarely in successful ones is a candidate cause. A difference that shows up in both is probably not.

06

Quantify the impact, then rank

A list of fifty findings is not much better than a pile of traces. For each failure pattern, estimate how many runs it affects, which workflows those runs belong to, and what it costs in failed tasks, latency, and spend. Then fix the top few. Everything else can wait for the next pass.

07

Turn each confirmed failure into a regression check

Once a pattern is fixed, save representative traces as an eval case so it cannot quietly come back. Over time, your eval set stops being a guess about what might go wrong and becomes a record of what actually did.

Trace signals that point to specific failures

You do not need to read every trace to know where to look. These are the signals we find most often behind real failures, and the failure each one usually indicates.

Signal in the traceUsually means
A success message with no matching write or API callThe agent claimed it finished work it never did.
The same tool called repeatedly with the same argumentsA retry loop without a change in strategy, often caused by an unhelpful tool error.
A tool error followed by a confident answerThe agent ignored the failure and filled the gap with a guess.
Input tokens growing faster in failing runs than in successful onesContext bloat: oversized tool results or repeated history drowning the instructions. See context bloat.
The agent asking the user for something it could have looked upA missing tool, a poorly described tool, or missing context.
A sub-agent returning far more data than the parent usesA hand-off that adds tokens and latency without improving the answer.
A policy-relevant step missing before a writeA skipped verification or confirmation step, which is a safety issue as well as a quality one.

Be honest about what automation can do

It would be convenient if you could hand a long trace to a strong model and ask it what went wrong. The research says you cannot, at least not reliably.

The TRAIL benchmark gave long-context models 148 human-annotated agent traces and asked them to locate the errors. The best model, Gemini 2.5 Pro, scored 11%. The Who&When study of 127 multi-agent systems found that the best automated method identified the responsible agent 53.5% of the time, but pinpointed the failing step only 14.2% of the time. Even strong reasoning models did not reach practical usability.

That does not mean automatic detection is hopeless. It means a single model reading a single trace is the wrong unit of analysis. Detection gets much more reliable when you narrow the question (one workflow, one rubric item), let code answer what code can answer, and look for differences across hundreds of runs instead of trying to find the needle in one.

It also means detection costs something. Reading traces carefully, judging them against specific rubrics, and clustering cohorts takes more model calls than sampling a few runs with a generic prompt. The cheap version gives you a number. The thorough version gives you a cause. Papaya is deliberately built for the thorough version, spending the model calls needed to analyze each use case in depth.

Worked example: a travel support agent

From 10,000 traces to three fixes

Imagine a travel support agent with 10,000 conversations in a week. The overall resolution rate is 81%, which nobody considers alarming.

ClassifyThe traces split into six use cases. Bookings resolve at 93%, baggage questions at 90%, but refunds resolve at only 58%.
DetectDeterministic checks flag refund runs where the agent confirmed a refund without calling the refund tool. A judge flags runs that quoted the wrong refund window.
CompareFailing refund runs call the booking lookup tool, which returns the full itinerary history. Successful runs use the lighter fare-rules lookup. The failing runs carry roughly three times the context.
RankThe context issue affects most failed refunds, the missing refund call a smaller share, and the wrong policy quote the rest. Fix them in that order.

The numbers here are illustrative, but the shape is typical. The average hid a broken workflow, and the cause was a decision about which tool to call, not an error anyone would have been paged for.

Where Papaya fits

Papaya does this procedure for you. It reads your traces in whatever format you already produce, classifies every run into the business use case it served, and splits each use case into succeeding and failing cohorts. Its 200+ research-backed analyses then look for the patterns that separate them, from missing tool calls to context bloat to sub-agents returning too much data. You get the top issues ranked by quantified impact, not a dashboard to dig through, and with code access Papaya can open the pull request for the fix. See how it works.

Frequently asked questions

How do you detect AI agent failures automatically?

Classify production traces by the job each run was doing, label outcomes using signals like final state, user feedback, and escalations, run deterministic checks for tool and policy failures, use an LLM judge with a workflow-specific rubric for semantic failures, and compare failing runs with successful runs of the same job to find the decisions that separate them.

Why don’t AI agent failures show up in error rates?

Most agent failures are wrong decisions rather than exceptions. The model returns a well-formed answer and every tool call succeeds, but the agent skipped a step, called the wrong tool, ignored a tool error, or claimed work it never did. Only the trace shows the problem.

Can an LLM find the failure in an agent trace by itself?

Not reliably. On the TRAIL benchmark the best long-context model scored 11% at locating errors in agent traces, and on Who&When the best method pinpointed the failing step only 14.2% of the time. Narrow rubrics, deterministic checks, and comparisons across many runs are far more reliable than asking one model to read one trace.

Do I need labeled data to detect agent failures?

No. Start with outcome signals you already have, such as final database or ticket state, repeat requests, thumbs-down feedback, escalations, and reopened tickets. Where no signal exists, a workflow-specific judge can provide provisional labels that you calibrate against a small set of human reviews.

What is the most common AI agent failure in production?

It varies by agent, but the patterns we see most often are claiming success without doing the work, retry loops without a change in strategy, ignoring tool errors, context bloat from oversized tool results, and skipped verification steps. Our field guide to AI agent failure patterns covers each one.

Which tools detect AI agent failures automatically?

Papaya detects failures automatically per business use case, comparing succeeding and failing runs with 200+ research-backed analyses and ranking findings by impact. Other tools with automatic detection include LangSmith Engine, Arize Signal, Braintrust Topics, Laminar, Latitude, and Raindrop. They differ in whether detection is scheduled or on request, whether you write the detectors yourself, and whether they propose fixes.

Sources used in this article

Related