You can evaluate multi-step AI agents without labeled data by combining signals you already have: the final state the agent left behind, what users did next, rules every correct run must follow, narrow reference-free LLM judges, consistency across repeated attempts, and comparisons between groups of similar runs. None of these is a perfect label on its own. Together they tell you which runs failed and why, and they tell you exactly which small set of runs is worth labeling by hand.
This guide explains each signal, what it catches, where it misleads, and how to combine them into an evaluation you can trust.
Why agents rarely come with labels
Labeled datasets work well for classification or short answers, where a person can say quickly whether an output is right. Agent runs are different. A single run may involve dozens of model calls and tool calls, take minutes, and depend on external systems. Labeling it means reconstructing what the user wanted, what the agent did, and whether each step was reasonable. That is expensive, slow, and hard to keep consistent across reviewers.
Worse, labels go stale. Users start asking for new things, the agent changes, and last quarter’s labeled set stops covering this quarter’s traffic. So most teams need evaluation that works first, with labels added where they help most.
Six signals that work without labels
Final state
The strongest evidence is often in your own systems. Was the refund recorded? Did the booking change? Did the ticket close and stay closed? τ-bench evaluates agents by comparing the database state at the end of a conversation with the expected goal state, and you can apply the same idea to production with invariants: a refund exists only if the order was eligible, a changed booking matches the date the user asked for.
Misleads when the right outcome depends on what the user meant, not what the system recorded.
What the user did next
Users label runs constantly without meaning to. They ask the same question again, rephrase, abandon the conversation, ask for a human, give a thumbs-down, or open a support ticket the next day. Tie these signals back to the run that caused them.
Misleads when users give up quietly. Silence is not success, so treat these signals as one-sided: strong evidence of failure, weak evidence of success.
Trajectory rules
Many requirements can be written as rules about the path, not the answer. Eligibility must be checked before a refund. A confirmation must precede a write. The same tool should not be called five times with the same arguments. A success message must follow a successful state-changing call. Checking these needs no labels, only code.
Misleads when a rule is too strict and flags valid alternative paths. Keep rules to things that must always be true.
Narrow, reference-free LLM judges
A judge does not need a reference answer if you ask it a narrow question about evidence it can see. “Does the final message describe the itinerary the booking tool returned?” “Did the agent answer the question the user actually asked?” For long runs, an agentic judge that can inspect intermediate steps does better. In the Agent-as-a-Judge study, it substantially outperformed a plain LLM judge and was about as reliable as the human evaluation baseline.
Misleads when the question requires knowledge the judge does not have. JudgeBench found strong models barely beat random on objectively hard correctness questions.
Consistency across repeated attempts
Run the same task several times. If the agent reaches different final states, at least some of those runs are wrong, and you learned that without a label. τ-bench formalized this as pass^k, the chance that an agent succeeds on all k attempts, and found it falls fast: even strong agents solved fewer than 25% of retail tasks reliably across eight tries.
Misleads when the agent is consistently wrong. Agreement is not correctness.
Comparison between groups of runs
Group runs by the business use case they served, then use the signals above to split each group into likely successes and likely failures. Compare the two. Behaviors that show up far more often in the failing group, such as a skipped lookup, a different tool, a much longer context, or a sub-agent returning pages of data, are strong candidates for causes, even though no single run was labeled.
Misleads when the groups are mixed. Comparing refund runs with booking runs finds differences between workflows, not between successes and failures.
How to combine the signals
No single signal is reliable enough to act on alone, but they fail in different ways, which is what makes the combination useful.
| When the signals say | Treat the run as |
|---|---|
| State check fails, or a trajectory rule is broken | A failure. Code checks are the most trustworthy evidence you have. |
| State is correct, the user came back with the same request | A likely failure the state check missed. Worth a closer look. |
| Judge says fail, nothing else does | Uncertain. A good candidate for human review. |
| Repeated attempts disagree | A reliability problem in this workflow, regardless of which attempt was right. |
| All signals agree it succeeded | A likely success. Sample a few for review anyway. |
Combining these signals by hand is tedious, which is why Papaya does it automatically for every business use case: outcome signals, customer feedback, and workflow-specific judging decide which runs succeeded, and cohort comparison finds out why the others did not.
Spend your labeling budget where it counts
Once these signals are running, labeling stops being a project and becomes a targeted activity. Label the runs where signals disagree, where the judge is unsure, and where a new cluster of failures has just appeared. A few dozen labels chosen this way teach you more than hundreds chosen at random, and they double as a calibration set for your judges.
Expect your criteria to change as you label. Researchers studying evaluation tooling call this criteria drift: reviewing outputs is how people discover what they actually care about. Build the process so your rubrics can evolve with what you learn.
Where Papaya fits
Papaya is designed for exactly this situation. It reads your traces in Langfuse, LangSmith, Arize Phoenix, or Braintrust format, or through its own SDK, classifies every run into the business use case it served, and combines outcome signals, customer feedback, and workflow-specific judging to separate succeeding runs from failing ones. Its interactive LLM judge builds the rubric for each use case from your trace data and customer signals, and 200+ research-backed analyses compare the cohorts to find what is going wrong, ranked by impact. No labeled dataset required to start. See how it works.
Frequently asked questions
How do you evaluate an AI agent without labeled data?
Combine signals that need no labels: final-state checks in your own systems, what users did after the run, rules every correct trajectory must follow, narrow reference-free LLM judges, consistency across repeated attempts, and comparisons between likely-successful and likely-failing runs of the same workflow. Then label only the runs where those signals disagree.
What is reference-free evaluation?
Reference-free evaluation judges an output without a known correct answer to compare against. For agents, it works best with narrow questions about evidence in the trace, such as whether the final message matches what a tool returned, rather than open-ended quality scores.
What is pass^k?
pass^k, introduced with τ-bench, is the probability that an agent succeeds on all of k repeated attempts at the same task. It measures reliability rather than single-attempt accuracy, and it falls quickly for most agents: in τ-bench’s retail domain, pass^8 was below 25% even for strong function-calling agents.
Can implicit user feedback replace labels?
Partly. Repeat requests, abandonment, escalations, thumbs-down, and reopened tickets are strong evidence that a run failed, but their absence is weak evidence of success, because many users give up quietly. Use implicit feedback alongside state checks and judges rather than on its own.
How many labels do I need to evaluate an agent?
Fewer than you think if you choose them well. Label the runs where your automatic signals disagree or a judge is uncertain, and add new failure clusters as they appear. A few dozen targeted labels per workflow are usually enough to calibrate judges and confirm patterns.
Which tools can evaluate AI agents without labeled data?
Papaya evaluates agents without a labeled dataset by combining outcome signals, customer feedback, and an LLM judge whose rubric it builds from your traces, then comparing succeeding and failing runs per business use case. Most other platforms, including Langfuse, LangSmith, Braintrust, and Arize, support reference-free LLM judges that you configure yourself.
Sources used in this article
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains — state-based evaluation and the pass^k reliability metric.
- Agent-as-a-Judge: Evaluate Agents with Agents — agentic judges that inspect intermediate steps.
- JudgeBench: A Benchmark for Evaluating LLM-based Judges — the limits of LLM judges on objective correctness.
- Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences — criteria drift.
Related
Keep reading
Agent evaluation
LLM-as-a-Judge for AI Agents: How to Build a Judge You Can Trust
LLM judges agree with people on preferences and struggle with correctness. How to build an agent judge from narrow, workflow-specific questions and calibrate it before you trust it.
Agent evaluation
How to Evaluate AI Agents in Production: A Practical Framework
In one public benchmark we analyzed, 22% of runs confidently told the customer the job was done without ever calling the real API. A guide to evaluating agents beyond the pass rate.

