The failure patterns that show up most in production AI agent traces are: claiming a task is done without doing it, retrying the same failing step without changing anything, ignoring tool errors, calling the wrong tool or passing the wrong arguments, drowning the model in context, missing the context it actually needed, skipping verification, not knowing when to stop, acting on stale state, and writing without confirmation. Almost none of them raise an error. All of them are visible in the trace if you know what to look for.
This guide describes each pattern the way it appears in real traces: what it looks like, why it happens, how to detect it automatically, and what usually fixes it.
Why patterns matter more than individual failures
One bad trace is an anecdote. You can stare at it for an hour, find something odd, and change a prompt, only to discover the odd thing appears just as often in runs that succeeded.
A pattern is different. It is a behavior that repeats across many runs of the same workflow and shows up far more often when the run fails. That is what makes it worth fixing. It also tells you how much a fix is worth, because you can count the runs it affects.
The research points the same way. The MAST study of multi-agent systems annotated 1,642 execution traces across seven frameworks and grouped what went wrong into 14 failure modes in three categories: system design issues, inter-agent misalignment, and task verification. One of its most useful findings is that failure modes are not equally fatal. Some, like being unaware of termination conditions, appear almost exclusively in failed runs. Others, like missing or incorrect verification, show up frequently even in runs that succeeded. If you only count how often a behavior occurs, you will fix the wrong things first.
Ten failure patterns, as they appear in traces
False completion: “done” without doing it
The agent tells the user the booking was changed, the refund was issued, or the ticket was updated, but no corresponding write or API call appears in the trace. In our analysis of the public τ²-bench airline benchmark, 22% of runs did exactly this. The final message looks perfect, so nothing downstream notices.
Detect a success claim in the final message with no matching state-changing tool call earlier in the run.
Usual fix require the tool result before the confirmation step, and make the confirmation quote the result.
Retry loops without a change in strategy
The same tool is called three, five, or ten times with identical or near-identical arguments. Each attempt fails the same way. The agent either gives up or eventually returns something vague. These loops are expensive and slow, and they usually trace back to a tool error message the model cannot act on.
Detect repeated calls to one tool with unchanged inputs inside a single run.
Usual fix make tool errors say what to do differently, and cap retries with an explicit fallback or escalation.
Ignored tool errors
A tool returns an error or an empty result, and the next message proceeds as if it had succeeded. The model fills the gap with something plausible. This is one of the most direct routes from a tool problem to a hallucination.
Detect an error or empty tool response followed by a confident answer that depends on it.
Usual fix return structured errors the model cannot mistake for data, and instruct the agent to say when it could not complete a step.
Right tool, wrong arguments, or the wrong tool entirely
The agent searches by name when it should search by booking reference, passes a date in the wrong format, or uses a general search tool when a specific lookup exists. The run often still finishes, just with the wrong data.
Detect schema or validation failures, and tool choices that differ between failing and successful runs of the same workflow.
Usual fix sharpen tool names and descriptions, constrain argument formats, and remove overlapping tools.
Context bloat
A tool or sub-agent returns far more than the next step needs: a full itinerary history, an entire document, a raw API payload. Input tokens climb with every turn, latency follows, and the instructions that matter get buried. We cover this in detail in context bloat.
Detect input tokens per step growing faster in failing runs than in successful ones; tool output that is never referenced later.
Usual fix return only the fields the next step uses, summarize sub-agent results, and stop replaying old tool output.
Missing context
The opposite problem. The information the model needed was neither in the prompt nor reachable through a tool, so it asks the user for something the system already knows, or guesses. Users experience this as an agent that keeps asking questions.
Detect clarifying questions for data the system holds; answers that contradict records the agent never retrieved.
Usual fix add the lookup, or pass the field in when the conversation starts.
Skipped or ineffective verification
The workflow includes a check, such as confirming eligibility, re-reading the final state, or validating a calculation, but the agent skips it, or runs it and ignores the result. MAST found verification problems even in successful runs, which is exactly why they are easy to overlook until a run fails because of them.
Detect required verification calls that are missing, or whose result does not affect the next step.
Usual fix move verification into code where possible, and make the agent’s next action depend on it.
Not knowing when to stop
Some runs end too early, before the task is complete. Others keep going long after the useful work is done, re-checking, re-summarizing, or asking whether there is anything else. MAST lists both premature termination and unawareness of termination conditions as distinct failure modes, and the second is one of the most strongly associated with failure.
Detect step counts far outside the workflow’s normal range; runs that end without reaching the required state.
Usual fix define explicit completion criteria per workflow, and check them in code.
Decisions on stale state
The agent reads a value early in the run, the world changes (a seat is taken, a price updates, the user corrects themselves), and the agent acts on the old value anyway. These failures are intermittent, which makes them hard to reproduce by hand.
Detect a write that uses a value from an earlier read when a later observation changed it.
Usual fix re-read critical state immediately before writes, or make writes conditional on the expected value.
Unsafe writes
The agent changes customer data, spends money, or makes an irreversible decision without confirmation, idempotency, or a way to roll back. Even when the outcome is correct, this is a risk you want to see before a customer does.
Detect state-changing calls with no preceding confirmation or policy check; duplicate writes from retries.
Usual fix require confirmation for risky actions, add idempotency keys, and put policy limits in code.
The same pattern means different things in different workflows
A retry loop in a search workflow costs a few seconds. The same loop in a payment workflow can charge a customer twice. Context bloat might be harmless in a short FAQ conversation and fatal in a long multi-step booking change.
That is why patterns need to be measured per business use case, not across the whole agent. A support agent that handles bookings, cancellations, baggage, insurance, and refunds is really five or six different workflows sharing a model. Each one has its own failure profile and needs its own fixes. Averaging them together hides the workflow that is quietly failing. Papaya measures every pattern this way: per business use case, against successful runs of the same workflow.
Quick reference
| Pattern | Trace signature | First fix to try |
|---|---|---|
| False completion | Success message, no matching write | Gate confirmation on the tool result |
| Retry loop | Same call, same arguments, repeated | Actionable tool errors; retry cap |
| Ignored tool error | Error or empty result, then a confident answer | Structured errors; permission to say “I couldn’t” |
| Wrong tool or arguments | Validation failures; tool choice differs from successful runs | Clearer tool descriptions; constrained inputs |
| Context bloat | Input tokens grow faster in failing runs | Trim tool output; summarize sub-agents |
| Missing context | Asks the user for known data | Add the lookup or pass the field |
| Skipped verification | Required check missing or ignored | Move the check into code |
| Not knowing when to stop | Step count outliers; required state missing | Explicit completion criteria |
| Stale state | Write uses an outdated read | Re-read before writing |
| Unsafe write | Write without confirmation or idempotency | Confirmation, idempotency keys, policy in code |
If you want the procedure for finding these automatically across thousands of traces, read how to detect AI agent failures automatically from traces.
Where Papaya fits
Papaya looks for all of these patterns, and many more, without asking you to dig through JSON. It classifies every trace into the business use case it served, compares succeeding and failing runs within each one, and runs 200+ research-backed analyses to find the patterns that actually separate them. Each finding comes with the number of runs it affects and its impact on quality, latency, and cost, so you know which one to fix first. See the research behind the checks.
Frequently asked questions
What are the most common AI agent failure patterns?
The most common production patterns are false completion, retry loops without a change in strategy, ignored tool errors, wrong tool or wrong arguments, context bloat, missing context, skipped verification, not knowing when to stop, decisions on stale state, and unsafe writes. Most of them do not raise errors and are only visible in the trace.
Why do AI agents fail in production?
Usually because of decisions rather than crashes. The agent picks the wrong tool, misreads a tool result, carries too much or too little context, skips a verification step, or stops at the wrong time. Real traffic also mixes many different user goals, and each workflow fails in its own way.
What is the MAST failure taxonomy?
MAST is a taxonomy of multi-agent system failures from the paper “Why Do Multi-Agent LLM Systems Fail?” It was built from expert annotation of agent traces and groups 14 failure modes into three categories: system design issues, inter-agent misalignment, and task verification. The accompanying dataset covers 1,642 annotated traces from seven frameworks.
How do I know which failure pattern to fix first?
Compare failing and successful runs of the same workflow. Prioritize patterns that appear far more often in failing runs, affect many runs, and sit in workflows that matter to your business. A pattern that is equally common in successful runs is rarely the cause.
Are AI agent hallucinations a failure pattern?
Hallucination is usually a symptom of one of the patterns above rather than a separate pattern. The most common causes in agent traces are ignored tool errors, missing context, and false completion, where the agent fills a gap with something plausible instead of saying it could not complete a step.
Is there a tool that finds these failure patterns automatically?
Yes. Papaya looks for all of these patterns automatically in production traces, per business use case, and ranks them by how many runs they affect and their impact on quality, latency, and cost. LangSmith Engine, Arize Signal, Braintrust Topics, Laminar, Latitude, and Raindrop also surface recurring failures, with different levels of automation.
Sources used in this article
- Why Do Multi-Agent LLM Systems Fail? — the MAST taxonomy: 14 failure modes in three categories, 1,642 annotated traces, and the finding that some failure modes appear almost exclusively in failed runs while verification failures also appear in successful ones.
- τ²-Bench: Evaluating Conversational Agents in a Dual-Control Environment — the public airline benchmark behind our 22% false-completion finding.
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains — evidence that agents are inconsistent across repeated attempts, which is why patterns must be measured across many runs.
Related
Keep reading
Agent reliability
How to Detect AI Agent Failures Automatically from Traces
Most agent failures never throw an error. A practical procedure for finding them in production traces, and an honest look at what automated detection can and cannot do.
Context engineering
Context Bloat: Too Much Context Makes Agents Worse
Agent prompts grow every turn: old messages, raw tool results, documents fetched just in case. This buildup makes agents slower, costlier, and less accurate. Here is how to trim it without losing information the agent still needs.

