To turn agent traces into improvements, run them through a loop: group runs by the job the user wanted done, label which ones succeeded, find the patterns that separate failing runs from successful ones, estimate what fixing each pattern is worth, ship the most valuable fix, and check that it actually moved the numbers on real traffic. Then do it again. Most teams collect traces faithfully and stall at the second step, because reading them by hand does not scale and dashboards do not tell you what to change.
This guide describes each stage of that loop, how to prioritize what comes out of it, and what kinds of changes it usually produces.
Why traces pile up without making the agent better
Instrumenting an agent is now easy. Every major framework and observability tool can capture model calls, tool calls, and sub-agents. The result, for most teams, is a growing archive that someone opens when a customer complains.
The problem is not a lack of data. It is that turning data into a change requires a chain of judgments: which runs failed, why, whether it matters, and what to change. Each link takes time from the engineers who also have to build the product. So the chain breaks, and the traces sit there.
Dashboards do not close the gap. An average success rate tells you something is wrong, not where. A token chart tells you costs went up, not which step caused it. Improvement needs a cause and a change, not a metric. Papaya was built to close exactly this gap, by running the loop below for you.
The agent improvement loop
Segment runs by business use case
A production agent does many jobs. A travel support agent books, cancels, rebooks, refunds, and answers baggage and insurance questions. Each job succeeds for different reasons, so each needs its own analysis. Classify every run by the use case it served before you measure anything. Otherwise a broken workflow that handles 8% of traffic disappears into an average that looks fine.
Label outcomes
For each use case, decide what success means and label runs with the best evidence you have: final state in your database, whether the user had to ask again, feedback, escalations, reopened tickets. Where no signal exists, a workflow-specific LLM judge can provide a provisional label, calibrated against a small set of human reviews.
Find what separates failing runs from successful ones
Within each use case, compare the failing cohort with the successful one. Cluster inside each cohort and look for differences in behavior: a tool that failing runs call and successful runs do not, a verification step that is skipped, a sub-agent that returns far more data, a context that grows much faster. A difference that appears mostly in failing runs is a candidate cause. This is where most of the analytical work lives.
Quantify the impact of each finding
For each candidate cause, estimate how many runs it affects, in which use cases, and what it costs in failed tasks, latency, and spend. This turns a long list of observations into a short, ordered list of decisions.
Ship the change as code
Turn the top finding into a concrete change: a prompt edit, a tool description, a trimmed tool output, a new lookup, a verification step, a different model for one step. Make it a reviewed pull request, not an edit in a playground that nobody can trace back later.
Verify on real traffic and keep a regression check
After the change ships, compare the affected use case before and after. Did the failure pattern shrink? Did anything else get worse? Save representative failing runs as an eval case so the problem cannot quietly return. Then go back to step 1, because traffic, models, and users keep changing.
How to decide what to fix first
Not every finding deserves a fix. A simple way to rank them is runs affected × how much those runs matter × how much a fix is likely to help. Here is an illustrative example for the travel agent above.
| Finding | Why it ranks where it does |
|---|---|
| Refund runs load the full itinerary history and lose the refund policy in the noise | Affects most failed refunds, refunds are high-value, and trimming a tool output is a small, safe change. Fix first. |
| Booking changes sometimes confirm before the change call returns | Fewer runs, but a false confirmation is costly and erodes trust. Fix second. |
| Baggage answers use a larger model than they need | Many runs, low risk, saves latency and cost without affecting quality. A good third change. |
| Occasional odd phrasing in greetings | Visible but harmless. Leave it. |
What improvements usually look like
The changes that come out of this loop are rarely dramatic. Most are small, specific edits to one layer of the agent.
| Layer | Typical finding and change |
|---|---|
| Context | Fields the model needs are missing, or buried under content it never uses. Add the field, trim the rest. |
| Tools | Tools return too much, return unclear errors, or overlap with each other. Narrow outputs, make errors actionable, merge or rename tools. |
| Instructions | Conflicting or missing rules for a specific use case. Make the rule explicit where that use case is handled. |
| Verification | Required checks are skipped or ignored. Move the check into code and make the next step depend on it. |
| Workflow structure | Sub-agents that add latency without improving results, or steps in an order that forces rework. Collapse, reorder, or parallelize. |
| Model choice | A large model on an easy step, or a small one on a hard step. Route by step difficulty. |
From a project to a loop
Run this once and you will find a handful of valuable fixes. Run it continuously and the agent keeps getting better as traffic changes, because every fix and every new run feeds the next analysis. The end state is an agent that improves itself under human supervision: patterns are found automatically, fixes arrive as pull requests, people approve them, and the results are measured. Some teams eventually let low-risk fixes merge on their own.
Where Papaya fits
Papaya runs this loop for you. It reads your traces in Langfuse, LangSmith, Arize Phoenix, or Braintrust format, or through its own SDK, classifies every run into the business use case it served, and compares succeeding and failing runs with 200+ research-backed analyses. You get the top fixes ranked by their quantified impact on quality, latency, and cost, not a dashboard to dig through. Connect your code and Papaya opens the pull request; auto-approve the ones you trust and the agent starts healing itself. See how it works.
Frequently asked questions
How do you turn LLM traces into actionable improvements?
Group runs by the business use case they served, label which ones succeeded, compare failing runs with successful ones to find the behaviors that separate them, estimate the impact of each finding, ship the most valuable fix as a reviewed code change, and verify it on real traffic before keeping the failing runs as a regression check.
What is an agent improvement loop?
An agent improvement loop is a repeating process that turns production traces into changes: find recurring failures, quantify them, fix the most valuable ones, measure the result, and feed the outcome back into the next round of analysis. Run continuously, it lets an agent keep improving as traffic and models change.
How do I improve AI agent performance in production?
Start from evidence rather than intuition. Find the workflows where failures concentrate, identify what failing runs do differently from successful ones, and make small, specific changes to context, tools, instructions, verification, workflow structure, or model choice. Measure each change on real traffic for the affected use case.
What is a self-healing AI agent?
A self-healing agent is one whose failures are detected and fixed through an automated loop: production traces are analyzed for recurring problems, fixes are generated as code changes, and approved fixes are deployed and measured. Human review remains in the loop, and teams typically allow automatic merging only for low-risk changes.
How often should I analyze agent traces?
Continuously if you can, and at least after every significant change to prompts, tools, models, or traffic. Agent behavior drifts as users and models change, so a one-time analysis goes stale quickly.
Which tools turn agent traces into improvements automatically?
Papaya runs the full loop: it classifies traces by business use case, finds what separates failing runs from successful ones with 200+ research-backed analyses, ranks fixes by impact, and opens pull requests. LangSmith Engine, Arize Signal, and Latitude also propose fixes for recurring issues, while most observability tools leave the analysis to you.
Sources used in this article
- Why Do Multi-Agent LLM Systems Fail? — evidence that some failure modes concentrate in failed runs while others also appear in successful ones, which is why cohort comparison matters.
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains — agent inconsistency across repeated attempts, and why fixes must be measured across many runs.
- Papaya research — the optimization patterns behind Papaya’s analyses, from context economy to tool quality and verification.
Related
Keep reading
Article
Agent Optimization vs. Observability: Why Watching Isn't Fixing
Your dashboard saw the bad run and did nothing. The gap isn't between having data and not — it's between seeing what happened and knowing what to change.
Agent reliability
How to Detect AI Agent Failures Automatically from Traces
Most agent failures never throw an error. A practical procedure for finding them in production traces, and an honest look at what automated detection can and cannot do.

