Skip to content

How to evaluate an AI agent without fooling yourself

By SunnyKumar Jonwal 10 min read

The most common way to ship a bad agent is to try it on five inputs, watch it succeed, and call it ready. Agents are nondeterministic, they take many steps, and they fail in odd ways on inputs you didn't think to try. Judging one by vibes is how you end up with a system that impresses in demos and disappoints at volume.

An evaluation suite fixes that. It doesn't need to be elaborate. A few dozen realistic tasks, a clear way to grade them, and the habit of running them on every change will put you ahead of most projects. This post shows how to build one and how to avoid the ways it can mislead you.

What you're trying to learn

Before writing any tests, decide which questions the evaluation should answer. The headline one is whether the agent completes the task, meaning what fraction of realistic tasks it gets right. Right behind it comes reliability: the same task run five times might succeed three of them, and that gap matters. You'll also want to know what it costs in tokens, dollars, and time, how it fails (wrong answer, crash, infinite loop, giving up, or doing something unsafe), and whether the last change helped or hurt. That last one is the daily use of an eval suite: regression detection.

Write these down. An eval that isn't tied to a decision is just a dashboard.

Start with tasks, not metrics

The foundation is a set of realistic tasks with a defined notion of success. Building it is mostly unglamorous work, and it's the part with the highest payoff.

Source them from reality. The best tasks come from real user requests, support tickets, or logs. If you don't have any yet, write what you'd expect users to ask, then replace them with real ones as soon as you can.

Cover the range. Include easy and hard cases, ambiguous requests, missing information, and tasks the agent should refuse or escalate. If everything in your set is a happy path, your score will be misleadingly high.

Start small. Twenty to fifty well-chosen tasks beat five hundred sloppy ones. Every task should have an owner who understands why it's there.

Define success before you run. For each task, write down what a correct result looks like. If you decide after seeing the output, you'll unconsciously grade on a curve.

Add failures as you find them. Every bug report becomes a new task. Over time, the suite becomes a memory of everything that has gone wrong.

Three levels of testing

Evaluate at more than one level, because problems hide at different depths.

  1. Unit tests for tools. Tools are ordinary code, so test them like it. Does search_orders return the right rows? Does safe_path reject ../ escapes? Do error messages say something useful? These tests are fast, deterministic, and catch a lot. See designing tools for LLM agents.

  2. End-to-end task tests. Give the agent a task in a realistic environment and check the outcome. This is the core of the suite and the source of your main success rate.

  3. Behavior and safety tests. Does it stay within its permissions? Does it refuse dangerous requests? Does it resist prompt injection embedded in a document it reads? These often matter more than raw accuracy, and they deserve their own tasks. Prompt injection and the lethal trifecta suggests what to include.

Grade the outcome, not the path

There are usually many valid routes to a correct result. An agent might read files in a different order than you'd expect, or use a search you didn't anticipate. If you grade on following a specific sequence of steps, you'll penalize good solutions and reward brittle ones.

Grade what the agent produced: the file it changed, the record it created, the answer it gave. For a coding task, that's whether the tests pass. For a data task, whether the output table matches. For a support task, whether the reply contains the right refund amount.

Some process checks are still worth having, but as guardrails and not as the score: it didn't call a forbidden tool, it stayed under a step limit, it didn't touch files outside its sandbox.

Ways to grade

Code-based graders. The gold standard when possible. Run tests, compare structured output to an expected value, check that a record exists. They're fast, cheap, repeatable, and objective. Use them wherever the task allows.

LLM-as-judge. For open-ended output like summaries, explanations, and drafted emails, a model can grade against a rubric. It scales well and it's surprisingly good, but it has known weaknesses:

  • Position bias. When comparing two answers, judges often favor whichever comes first.
  • Verbosity bias. Longer answers tend to score higher regardless of quality.
  • Self-preference. A model may rate its own family's outputs more generously.
  • Rubric drift. Vague criteria produce inconsistent scores.

Reduce these problems by writing specific rubrics ("mentions the 30-day window; cites the correct policy; doesn't promise a refund"), asking the judge to explain before scoring, randomizing order in comparisons, and using a different model as judge when possible. Then calibrate: have a person grade a sample, compare, and adjust until the judge agrees with the human at a rate you trust.

Human review. Slow and expensive, and irreplaceable for judging quality, tone, and edge cases. Use it to calibrate your automated graders and to review a rotating sample of real traffic, not to grade everything.

Nondeterminism: run it more than once

Run the same task twice and you may get two different paths and two different outcomes. That means a single run tells you very little.

Run each task several times, say five, and look at the distribution.

import statistics

def evaluate(tasks, run_agent, grade, trials=5):
    results = []
    for task in tasks:
        outcomes = []
        for _ in range(trials):
            run = run_agent(task["input"])          # returns answer, steps, tokens, seconds
            outcomes.append({
                "passed": grade(task, run),
                "steps": run["steps"],
                "tokens": run["tokens"],
                "seconds": run["seconds"],
            })
        rate = sum(o["passed"] for o in outcomes) / trials
        results.append({
            "task": task["id"],
            "pass_rate": rate,
            "always": rate == 1.0,
            "median_tokens": statistics.median(o["tokens"] for o in outcomes),
        })
    return results

summary = evaluate(TASKS, run_agent, grade)
print("mean pass rate:", statistics.mean(r["pass_rate"] for r in summary))
print("reliable (5/5):", sum(r["always"] for r in summary), "of", len(summary))

Two numbers are worth tracking. The mean pass rate says how often it succeeds on average. The fully reliable fraction, the tasks that pass every time, tells you how much you can depend on it without a human checking. For production use, the second number often matters more. A task that passes 60 percent of the time is a task you can't hand off.

Control the environment

Flaky evals are worthless, so remove randomness that isn't the model's.

  • Sandbox the environment so every run starts from the same state: a fresh copy of the repo, a seeded database, a clean folder.
  • Mock external services or use recorded responses, so a third-party outage doesn't fail your tests.
  • Freeze the clock and any random seeds where you can.
  • Reset between runs. State left over from a previous run can make the next one pass or fail for the wrong reasons.
  • Record everything: full transcripts, tool calls, and environment state, so you can replay and investigate failures.

Also pin the model version, prompt, and tool definitions for each run, and record them with the results. Without that, you can't tell whether a score changed because of your edit or because something else moved.

Look at more than success rate

Task success is the headline, and a few other numbers round out the picture. Cost per task, in tokens and dollars, matters because an agent that succeeds at ten times the cost may not be an improvement. Steps per task is a good confusion detector, since a sudden increase often signals looping. Latency is what users feel, and a high tool error rate points at bad descriptions or schemas. Safety violations, meaning any attempt to leave the sandbox, call a forbidden tool, or leak data, deserve their own count. The target is zero, and even one is worth a look. If refusal and escalation are meant to happen, track those rates too.

Plot them over time. A dip in success with a spike in steps tells a different story from a dip with flat steps.

Read the transcripts

Numbers say whether something is wrong. Transcripts say why. After every eval run, read the failures, not all of them if there are many, but a solid sample from each failure category.

Ask specific questions. At the step where it went wrong, did the model have the information it needed? Was a tool result confusing? Did it misread the task, or did the task description mislead it? Was the grader wrong, marking a valid answer as a failure?

That last one happens more often than people expect. When a "failure" looks right to you, the fix is in the grader or the task definition, and you should fix it rather than trying to make the agent match a flawed check.

Sort failures into categories: bad tool choice, bad arguments, misunderstood goal, context lost, gave up early, looped. The biggest category tells you where to work.

Common traps

Overfitting to your test set. If you tune prompts against the same 30 tasks repeatedly, scores rise while real-world performance doesn't. Keep a held-out set you look at rarely, and refresh the main set with new real cases.

Saturated evals. When everything passes, the suite has stopped telling you anything. Add harder tasks.

Leaked answers. If reference solutions or hints appear anywhere the agent can read, you're measuring cheating.

Unrepresentative tasks. A suite written entirely by the developers who built the agent tends to test what they thought of. Real users surprise you.

Trusting a single number. A 90 percent average can hide a category that fails every time.

Ignoring cost. A change that raises accuracy two points and triples spend may not be a win.

Grading by feel. If "good enough" isn't defined, every run will feel a little different.

Offline, then online

Offline evaluation, the suite you run on your own tasks, tells you whether a change is safe to ship. It can't tell you everything, because production traffic is messier.

Once live, add online monitoring: log real runs, sample them for review, track success proxies such as user thumbs-up, follow-up corrections, and escalation rates, and watch cost and latency. When you find a real failure, turn it into a new offline task. That loop, production failures feeding the test suite, is what makes a system steadily better instead of just different.

Run the suite in CI so a prompt tweak or dependency bump can't silently break behavior. Even a small suite, run on each pull request, catches regressions before users do. Dedicated tools exist to help with running evals and tracking results, including open-source options such as promptfoo and Inspect, but a script like the one above is a fine start.

A starter plan

  1. Write 25 realistic tasks with defined success criteria.
  2. Build a sandbox that resets to a known state.
  3. Write code-based graders where you can, and a rubric-based judge where you can't.
  4. Run each task five times and record pass rate, steps, tokens, and time.
  5. Read the failures and sort them into categories.
  6. Fix the biggest category, rerun, and compare.
  7. Add each newly discovered failure as a task.
  8. Put the suite in CI, and review a sample of real traffic weekly.

You'll be surprised how much a modest suite teaches you in the first week, and how much you'd been guessing before. Evals also make model upgrades far less scary: run the suite on the new model, compare, and decide with data. If you want to go deeper on designing what the agent sees, context engineering and tool design are the usual levers to pull once the numbers show you where it hurts.