To debug an AI agent in production, start from the bad outcome and read the trace backward until you find the first wrong decision. Then ask what the model could see at that moment, whether a tool misled it, and whether the same mistake happens in other runs of the same workflow. Reproduce the step several times, because agents are not deterministic, and judge your fix against a cohort of runs rather than the one that failed. Debugging an agent is less like reading a stack trace and more like reviewing a decision.
This guide walks through that process step by step, with the root causes we see most often and the places to look for each.
Why debugging agents is different
Traditional debugging assumes that a failure leaves evidence: an exception, an error code, a line number. Agents rarely fail that way. The model returns well-formed text, every tool call succeeds, and the run ends on time. The failure is that the agent chose the wrong tool, misread a result, skipped a step, or told the user something that was not true.
Two more things make it harder. Agents are non-deterministic, so the same input can succeed on one run and fail on the next. The τ-bench authors found that even strong function-calling agents succeeded on fewer than half of their tasks, and that the share of tasks solved reliably across eight attempts dropped below 25% in the retail domain. And production traffic mixes many different jobs, so a failure in one workflow can look like noise in the average.
A step-by-step debugging process
Connect the complaint to the run
Most production debugging starts with a signal outside the trace: a support ticket, a thumbs-down, an escalation, a customer who asked the same question twice. If you cannot find the exact run behind that signal, nothing else in this process works. Make sure every run carries a conversation or session ID, a user or account ID, and the agent and prompt version that handled it.
Read the trace backward and find the first wrong decision
Start at the bad outcome and walk back through the steps. At each one, ask whether the agent’s decision was reasonable given what it had seen so far. The first step where the answer is no is your real starting point. Everything after it is usually a consequence, and fixing a later step only hides the problem.
Check what the model could see at that step
Most wrong decisions are reasonable responses to bad context. Look at the exact input to that model call. Was the information it needed missing? Was it there but buried under pages of tool output or old history? Was it stale, from before the user corrected themselves? Were two instructions in conflict? Context problems are among the most common root causes, and they are invisible if you only look at outputs. See context bloat for the oversized-context case.
Check the tool layer
Look at the tool calls just before the wrong decision. Did a tool return an error the agent ignored? An empty result it treated as an answer? A huge payload that crowded out the instructions? Is the tool’s name or description close enough to another tool’s that the model picked the wrong one? Many agent bugs are really tool-design bugs.
Find out whether it is one run or a pattern
Search for runs of the same workflow with the same signature: the same tool error, the same missing call, the same kind of request. Then compare them with successful runs of that workflow. If the suspected cause appears in most failing runs and rarely in successful ones, you have found a pattern worth fixing. If it appears in both, keep looking. Papaya runs this comparison automatically for every business use case. We cover this in more detail in how to detect AI agent failures from traces.
Reproduce the step, more than once
Replay the failing step with the same inputs, several times. If it fails every time, you have a deterministic problem in context or tools. If it fails sometimes, you have a reliability problem, and your fix needs to raise the success rate, not just make one run pass. Replaying a single step is cheaper and more precise than rerunning the whole conversation.
Verify the fix on a cohort, then keep it as a regression test
A prompt change that fixes one run can break three others. Test the fix against a set of real runs from the same workflow, including ones that were already succeeding. Once it ships, keep representative failing runs as an eval case so the bug cannot quietly return.
Common symptoms and where to look
| Symptom | Likely cause and where to look |
|---|---|
| Agent says it did something it did not | False completion. Look for a missing state-changing tool call before the confirmation message. |
| Agent repeats the same action | A retry loop. Check the tool error message; the model usually cannot tell what to change. |
| Agent makes up a fact | Missing context or an ignored tool error. Check whether the fact was ever retrieved. |
| Agent asks the user for known information | A missing lookup tool or context that was never passed in. |
| Agent gets slower and worse over a long conversation | Context growth. Compare input tokens per step in failing and successful runs. |
| Agent uses the wrong tool | Overlapping or vague tool descriptions. Compare tool choice with successful runs. |
| Works in testing, fails in production | Real requests differ from test cases. Classify production traffic by use case and look for workflows your tests never covered. |
| Fails only sometimes | Non-determinism near a decision boundary. Replay the step several times and tighten the instruction or the tool contract. |
For a fuller list, see our field guide to AI agent failure patterns.
Debugging multi-agent systems
With several agents, the hardest question is which agent made the decisive mistake, and when. Researchers studying this found it genuinely hard to automate. On the Who&When dataset of 127 multi-agent systems, the best method identified the responsible agent 53.5% of the time but the decisive step only 14.2% of the time. Hand-offs are the usual suspects: one agent passes along too little context, or far too much, or a sub-agent redoes work the parent already did. Log every hand-off explicitly, with what was passed and why, so you can see exactly where the information changed.
Instrumentation that makes debugging possible
You cannot debug what you did not record. At minimum, capture for every run: the full input to each model call, every tool call with its arguments and result, sub-agent hand-offs, retries, the final state, and identifiers that link the run to users and to the agent version. OpenTelemetry’s generative AI conventions are a sensible default so you are not locked into one vendor. Be deliberate about personal data in traces, since you will be reading them and possibly sending them to an evaluation model.
Where Papaya fits
Debugging one run by hand is fine. Debugging thousands is not. Papaya reads your traces in Langfuse, LangSmith, Arize Phoenix, or Braintrust format, or through its own SDK, groups them by the business use case each run served, and compares failing runs with successful ones using 200+ research-backed analyses. Instead of you hunting for the first wrong decision, it shows you the recurring ones, how many runs each affects, and what fixing it is worth, and with code access it opens the pull request. See how it works.
Frequently asked questions
Why is my AI agent failing in production?
Usually because of a decision rather than a crash: the agent lacked the context it needed or had too much of it, misread or ignored a tool result, picked the wrong tool, or skipped a verification step. Production traffic also contains requests your tests never covered. Read the failing trace backward to find the first wrong decision, then check whether it recurs in other runs of the same workflow.
How do you debug an LLM agent?
Connect the complaint to the exact run, read the trace backward to find the first wrong decision, inspect what the model could see at that step and what the tools returned, check whether the same pattern appears in other runs, replay the step several times, and verify the fix on a cohort of real runs before keeping it as a regression test.
Why does my agent work in testing but fail in production?
Real users ask for things your test cases did not include, phrase them differently, and change their minds mid-conversation. Production also exposes long conversations, real tool latency and errors, and rare workflows. Classify production traffic by use case to find the workflows your tests never covered.
How do I reproduce a non-deterministic agent failure?
Replay the failing step with exactly the same inputs several times rather than rerunning the whole conversation. If it fails every time, the problem is in the context or tools. If it fails only sometimes, treat it as a reliability problem and measure your fix by how much it raises the success rate across repeated attempts.
What should I log to debug AI agents?
Log the full input to each model call, every tool call with arguments and results, sub-agent hand-offs, retries, the final state, and identifiers linking each run to the user, conversation, and agent version. OpenTelemetry’s generative AI conventions are a good vendor-neutral starting point.
What tools help debug AI agents in production?
You need tracing first, from tools such as Langfuse, LangSmith, Arize Phoenix, Braintrust, or Datadog. To find recurring causes across many runs, Papaya compares failing and successful runs per business use case and ranks what to fix by impact, and LangSmith Engine, Arize Signal, Braintrust Topics, Laminar, Latitude, and Raindrop also surface recurring issues.
Sources used in this article
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains — agents succeed on fewer than half of tasks and are inconsistent across repeated attempts.
- Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems — automated failure attribution across 127 multi-agent systems.
- Inside the LLM Call: GenAI Observability with OpenTelemetry — instrumentation for agent, model, and tool spans.
Related
Keep reading
Agent reliability
How to Detect AI Agent Failures Automatically from Traces
Most agent failures never throw an error. A practical procedure for finding them in production traces, and an honest look at what automated detection can and cannot do.
Agent reliability
AI Agent Failure Patterns: A Field Guide from Production Traces
False completion, retry loops, ignored tool errors, context bloat, skipped verification. The ten patterns behind most production agent failures, as they actually appear in traces.

