← Articles

Agent optimization · Practical guide

How to Turn Agent Traces into Improvements

Most teams collect traces and never act on them. The agent improvement loop that turns traces into ranked, shipped, and verified fixes, and what those fixes usually look like.

Macy Mody
· 7 minute read
Grey trace lines combed into blue strands that pass a golden marker and curve back into a continuous loop

To turn agent traces into improvements, run them through a loop: group runs by the job the user wanted done, label which ones succeeded, find the patterns that separate failing runs from successful ones, estimate what fixing each pattern is worth, ship the most valuable fix, and check that it actually moved the numbers on real traffic. Then do it again. Most teams collect traces faithfully and stall at the second step, because reading them by hand does not scale and dashboards do not tell you what to change.

This guide describes each stage of that loop, how to prioritize what comes out of it, and what kinds of changes it usually produces.

Why traces pile up without making the agent better

Instrumenting an agent is now easy. Every major framework and observability tool can capture model calls, tool calls, and sub-agents. The result, for most teams, is a growing archive that someone opens when a customer complains.

The problem is not a lack of data. It is that turning data into a change requires a chain of judgments: which runs failed, why, whether it matters, and what to change. Each link takes time from the engineers who also have to build the product. So the chain breaks, and the traces sit there.

Dashboards do not close the gap. An average success rate tells you something is wrong, not where. A token chart tells you costs went up, not which step caused it. Improvement needs a cause and a change, not a metric. Papaya was built to close exactly this gap, by running the loop below for you.

A trace you never act on costs you the storage. A pattern you find and fix pays for all of them.

The agent improvement loop

01

Segment runs by business use case

A production agent does many jobs. A travel support agent books, cancels, rebooks, refunds, and answers baggage and insurance questions. Each job succeeds for different reasons, so each needs its own analysis. Classify every run by the use case it served before you measure anything. Otherwise a broken workflow that handles 8% of traffic disappears into an average that looks fine.

02

Label outcomes

For each use case, decide what success means and label runs with the best evidence you have: final state in your database, whether the user had to ask again, feedback, escalations, reopened tickets. Where no signal exists, a workflow-specific LLM judge can provide a provisional label, calibrated against a small set of human reviews.

03

Find what separates failing runs from successful ones

Within each use case, compare the failing cohort with the successful one. Cluster inside each cohort and look for differences in behavior: a tool that failing runs call and successful runs do not, a verification step that is skipped, a sub-agent that returns far more data, a context that grows much faster. A difference that appears mostly in failing runs is a candidate cause. This is where most of the analytical work lives.

04

Quantify the impact of each finding

For each candidate cause, estimate how many runs it affects, in which use cases, and what it costs in failed tasks, latency, and spend. This turns a long list of observations into a short, ordered list of decisions.

05

Ship the change as code

Turn the top finding into a concrete change: a prompt edit, a tool description, a trimmed tool output, a new lookup, a verification step, a different model for one step. Make it a reviewed pull request, not an edit in a playground that nobody can trace back later.

06

Verify on real traffic and keep a regression check

After the change ships, compare the affected use case before and after. Did the failure pattern shrink? Did anything else get worse? Save representative failing runs as an eval case so the problem cannot quietly return. Then go back to step 1, because traffic, models, and users keep changing.

How to decide what to fix first

Not every finding deserves a fix. A simple way to rank them is runs affected × how much those runs matter × how much a fix is likely to help. Here is an illustrative example for the travel agent above.

FindingWhy it ranks where it does
Refund runs load the full itinerary history and lose the refund policy in the noiseAffects most failed refunds, refunds are high-value, and trimming a tool output is a small, safe change. Fix first.
Booking changes sometimes confirm before the change call returnsFewer runs, but a false confirmation is costly and erodes trust. Fix second.
Baggage answers use a larger model than they needMany runs, low risk, saves latency and cost without affecting quality. A good third change.
Occasional odd phrasing in greetingsVisible but harmless. Leave it.

What improvements usually look like

The changes that come out of this loop are rarely dramatic. Most are small, specific edits to one layer of the agent.

LayerTypical finding and change
ContextFields the model needs are missing, or buried under content it never uses. Add the field, trim the rest.
ToolsTools return too much, return unclear errors, or overlap with each other. Narrow outputs, make errors actionable, merge or rename tools.
InstructionsConflicting or missing rules for a specific use case. Make the rule explicit where that use case is handled.
VerificationRequired checks are skipped or ignored. Move the check into code and make the next step depend on it.
Workflow structureSub-agents that add latency without improving results, or steps in an order that forces rework. Collapse, reorder, or parallelize.
Model choiceA large model on an easy step, or a small one on a hard step. Route by step difficulty.

From a project to a loop

Run this once and you will find a handful of valuable fixes. Run it continuously and the agent keeps getting better as traffic changes, because every fix and every new run feeds the next analysis. The end state is an agent that improves itself under human supervision: patterns are found automatically, fixes arrive as pull requests, people approve them, and the results are measured. Some teams eventually let low-risk fixes merge on their own.

Where Papaya fits

Papaya runs this loop for you. It reads your traces in Langfuse, LangSmith, Arize Phoenix, or Braintrust format, or through its own SDK, classifies every run into the business use case it served, and compares succeeding and failing runs with 200+ research-backed analyses. You get the top fixes ranked by their quantified impact on quality, latency, and cost, not a dashboard to dig through. Connect your code and Papaya opens the pull request; auto-approve the ones you trust and the agent starts healing itself. See how it works.

Frequently asked questions

How do you turn LLM traces into actionable improvements?

Group runs by the business use case they served, label which ones succeeded, compare failing runs with successful ones to find the behaviors that separate them, estimate the impact of each finding, ship the most valuable fix as a reviewed code change, and verify it on real traffic before keeping the failing runs as a regression check.

What is an agent improvement loop?

An agent improvement loop is a repeating process that turns production traces into changes: find recurring failures, quantify them, fix the most valuable ones, measure the result, and feed the outcome back into the next round of analysis. Run continuously, it lets an agent keep improving as traffic and models change.

How do I improve AI agent performance in production?

Start from evidence rather than intuition. Find the workflows where failures concentrate, identify what failing runs do differently from successful ones, and make small, specific changes to context, tools, instructions, verification, workflow structure, or model choice. Measure each change on real traffic for the affected use case.

What is a self-healing AI agent?

A self-healing agent is one whose failures are detected and fixed through an automated loop: production traces are analyzed for recurring problems, fixes are generated as code changes, and approved fixes are deployed and measured. Human review remains in the loop, and teams typically allow automatic merging only for low-risk changes.

How often should I analyze agent traces?

Continuously if you can, and at least after every significant change to prompts, tools, models, or traffic. Agent behavior drifts as users and models change, so a one-time analysis goes stale quickly.

Which tools turn agent traces into improvements automatically?

Papaya runs the full loop: it classifies traces by business use case, finds what separates failing runs from successful ones with 200+ research-backed analyses, ranks fixes by impact, and opens pull requests. LangSmith Engine, Arize Signal, and Latitude also propose fixes for recurring issues, while most observability tools leave the analysis to you.

Sources used in this article

Related