← Articles

Agent evaluation · Practical guide

LLM-as-a-Judge for AI Agents: How to Build a Judge You Can Trust

LLM judges agree with people on preferences and struggle with correctness. How to build an agent judge from narrow, workflow-specific questions and calibrate it before you trust it.

Niyaz Puzhikkunnath
· 9 minute read
Blue and orange trajectories passing through fine calibration arcs and a graduated scale around a golden point

LLM-as-a-judge means using a language model to grade another model’s output against a rubric. For AI agents, it works when you use it for the right questions: let code check everything code can check, give the judge one narrow, workflow-specific question at a time, show it the parts of the trace that matter rather than only the final answer, and calibrate it against human labels before you trust its numbers. The best LLM judge for your agent is not a particular product. It is a judge whose agreement with your own reviewers you have actually measured.

This guide covers what the research says about LLM judges, why agents make judging harder, and a step-by-step way to build a judge you can rely on.

What the research actually says about LLM judges

The case for LLM judges is real. In the original MT-Bench and Chatbot Arena study, strong judges like GPT-4 matched human preferences more than 80% of the time, about the same rate at which humans agree with each other. That is why nearly every evaluation platform now ships some form of model-graded evaluator.

The same paper is also the best catalog of what goes wrong. It documented position bias (preferring whichever answer is shown first), verbosity bias (preferring longer answers), self-enhancement bias (preferring answers from the same model), and weak reasoning on questions like math. Later work sharpened each point:

  • A systematic study of position bias across 15 judges and more than 150,000 evaluation instances found that the bias is not random. It varies by judge and task, and grows when the candidates are close in quality.
  • LLM evaluators can recognize their own generations, and the more strongly they recognize them, the more they prefer them.
  • JudgeBench tested judges on objectively right-or-wrong pairs in knowledge, reasoning, math, and coding. Many strong models, including GPT-4o, performed only slightly better than random guessing.

Put together: judges agree with people about preferences surprisingly well, and are surprisingly unreliable about correctness. Agent evaluation is mostly a question of correctness. Did the agent do the right thing?

An LLM judge is a measuring instrument. You would not trust a scale you never checked against a known weight.

Why judging agents is harder than judging answers

Most judge guidance was written for single responses. Agents add three problems.

The answer is not the work. An agent that says “your booking is changed” may never have called the booking API. A judge that sees only the final message will grade it as excellent. You have to judge the trajectory: which tools were called, with which arguments, whether errors were handled, whether required checks happened.

Traces are long. Dumping a full multi-step trace into a judge prompt invites the same failures you are trying to detect: lost details, distraction, and confident summaries of things that did not happen. On the TRAIL benchmark, the best long-context model located only 11% of the errors in agent traces.

“Correct” depends on the workflow. A good refund conversation and a good booking conversation succeed for different reasons. A single generic rubric for “helpfulness” will blur them together and grade neither well.

How to build an LLM judge you can trust

01

Take everything code can check away from the judge

Did the required tool get called? Did the final state match? Did the agent stay within policy limits? Did a write happen without confirmation? These are deterministic questions with deterministic answers. Code is cheaper, faster, and does not drift. Save the judge for questions that need language understanding.

02

Write the rubric per workflow, from real failures

Group your traces by the job the user wanted done, then write criteria for each job. The best source of criteria is your own production data: runs that users complained about, runs that escalated, runs your team flagged. A rubric written in the abstract measures what you imagined could go wrong. A rubric written from traces measures what actually does.

Expect the rubric to change as you read more outputs. Researchers studying evaluation tools call this criteria drift: people need criteria to grade outputs, but grading outputs is also how they discover the criteria. Plan for revisions.

03

Ask one narrow question per criterion

“Rate this conversation from 1 to 10” produces scores nobody can act on. “Did the agent state the fare difference before changing the booking? Yes, no, or cannot tell” produces a signal you can count, debug, and fix. Prefer binary or short ordinal answers, ask for the evidence behind each answer, and always allow the judge to say it cannot tell.

04

Give the judge the evidence, not the whole trace

For each criterion, pass the parts of the trace it depends on: the user request, the relevant tool calls and results, and the final message. Shorter, targeted context makes the judge more accurate and much cheaper. For very long or complex runs, consider a judge that can navigate the trace itself. The Agent-as-a-Judge work found that an agentic judge with access to intermediate steps dramatically outperformed a plain LLM judge and was about as reliable as the human evaluation baseline.

05

Calibrate against human labels before trusting the numbers

Have people label a sample of traces for each criterion, typically 50 to 100 per workflow to start, deliberately including hard and borderline cases. Run the judge on the same traces and measure agreement. Look beyond raw accuracy: if failures are rare, a judge that always says “pass” will look accurate. Check how many real failures it catches and how many of its failure calls are real.

Measure recall and precision on the failure class, Cohen’s kappa against human labels, and the specific cases where the judge and the humans disagree.

06

Control for known biases

For pairwise comparisons, run both orders and treat disagreement as a tie. Do not let length stand in for quality: tell the judge to ignore it, and check that scores do not simply rise with response length. Where possible, use a judge from a different model family than the agent it grades. Run the judge more than once on a sample to check that its answers are stable.

07

Re-check it as the agent and traffic change

A judge calibrated in March can quietly drift by June. The agent changed, users started asking different things, or the judge model was updated. Re-run your calibration set on a schedule and whenever a major change ships, and add new disagreements to it.

Worked example: a booking-change rubric

Workflow: a customer asks a travel agent to move their flight to a different date.

Code checksChange-booking tool called; final reservation matches the requested date; fee within policy; confirmation step present before the write.
Judge question 1Did the agent tell the customer the fare difference and any change fee before making the change? Yes / no / cannot tell.
Judge question 2Did the final message accurately describe the new itinerary as returned by the tool? Yes / no / cannot tell.
Judge question 3If the requested date was unavailable, did the agent offer the closest alternatives instead of failing silently? Yes / no / not applicable.

Each judge question gets only the evidence it needs: the customer’s request, the relevant tool results, and the agent’s messages. Each is calibrated separately against a few dozen human-labeled conversations. When the judge and the reviewers disagree, the disagreement usually reveals either an ambiguous criterion or a real gap in the rubric.

What judging costs, and when it is worth it

Careful judging is not free. Several narrow questions per run, each with its own evidence, cost more model calls than one generic score. But compare it with the alternative. In the Agent-as-a-Judge study, three human experts spent a self-reported 86.5 hours evaluating the benchmark, about $1,297.50 at the authors’ assumed rate. The agentic judge took about two hours and cost $30.58 in API calls.

The practical rule: sample broadly with cheap deterministic checks, judge deeply where the stakes are high or the outcome is unclear, and spend human time on calibration and on the cases where the judge is unsure.

Which tool should you use for LLM-as-a-judge?

Most evaluation and observability platforms now let you define LLM judges, including Langfuse, LangSmith, Braintrust, Arize Phoenix, Opik, and DeepEval. They differ in workflow and pricing more than in what a judge can do. For most teams the tool matters less than three questions: can you write per-workflow criteria, can you feed the judge the right slice of the trace, and can you measure the judge against human labels? If a tool makes any of those hard, you will end up with a dashboard of scores nobody trusts.

Papaya takes a different approach to the first question. Instead of asking you to write criteria, its interactive LLM judge builds the rubric for each business use case from your trace data and customer signals. You review it and tell it what to change in plain language. The judge’s verdicts then feed Papaya’s comparison of failing and successful runs, so the output is a ranked list of fixes rather than a column of scores.

Where Papaya fits

Papaya’s interactive LLM judge builds the evaluation rubric for you, from your trace data and customer signals, separately for each business use case your agent handles. You review it, tell it what to change in plain language, and it keeps code-checkable questions out of the judge. The judge’s verdicts then feed Papaya’s analysis of what separates failing runs from successful ones, so you get ranked fixes, not just scores. See how it works.

Frequently asked questions

What is LLM-as-a-judge?

LLM-as-a-judge is the practice of using a language model to evaluate another model’s output against a rubric or reference. It scales review that would otherwise need people, but it must be calibrated against human labels because judges have known biases and can be unreliable on questions of correctness.

What is the best LLM-as-a-judge tool for evaluating agents?

Most platforms, including Langfuse, LangSmith, Braintrust, Arize Phoenix, Opik, and DeepEval, support LLM judges that you configure yourself. Papaya’s interactive LLM judge builds the rubric for you from trace data and customer signals, per business use case. The deciding factors are whether you get workflow-specific criteria, whether the judge sees the relevant parts of the trace rather than only the final answer, and whether its agreement with human labels is measured.

How accurate are LLM judges?

On preference questions, strong judges agreed with humans more than 80% of the time in the MT-Bench study, similar to human-to-human agreement. On objective correctness, JudgeBench found many strong models performing only slightly better than random. Accuracy for your agent has to be measured on your own labeled traces.

What biases do LLM judges have?

The best documented are position bias, verbosity bias, and self-enhancement bias, where judges favor answers from their own model. Mitigate them by running pairwise comparisons in both orders, controlling for length, using a judge from a different model family, and checking repeated runs for consistency.

What is agent-as-a-judge?

Agent-as-a-judge uses an agentic system, one that can inspect intermediate steps and artifacts, to evaluate another agent. In the original study it substantially outperformed a plain LLM judge, was about as reliable as the human evaluation baseline, and cost a small fraction of human review.

How many labeled examples do I need to calibrate a judge?

Start with roughly 50 to 100 human-labeled traces per workflow, including hard and borderline cases, then grow the set with every disagreement you find. The goal is not a perfect number but a measured one: known agreement, known blind spots.

Sources used in this article

Related