Glossary · Agent architecture

Agent evaluation

Testing an AI agent on defined tasks and grading both the outcome and the steps it took, to measure what it can do and to catch regressions.

Agent evaluation is the practice of running an AI agent on defined tasks and grading the results, both the final outcome and the path the agent took, to measure what it can do and to catch regressions.

Vocabulary. Anthropic’s “Demystifying evals for AI agents” (January 2026) sets out clear terms. A task has inputs and success criteria. A trial is one attempt at a task, and because agents are not deterministic, each task runs several trials. Graders score aspects of the result. The transcript, also called a trace or trajectory, records every output and tool call in a trial. The outcome is the final state of the environment, such as whether a refund record actually exists.

What to grade.

  • Outcome: did the environment end up in the right state?
  • Final response: is the answer correct, complete and within policy?
  • Trajectory: did the agent call the right tools, with the right arguments, in a sensible order? Google ADK compares an agent’s actual tool calls with an expected trajectory, alongside its final response, using .test.json files for quick unit-style checks and evalsets for longer multi-turn sessions.

Anthropic advises grading what the agent produced over the exact path it took, since there is often more than one valid path.

Graders. Code-based graders are fast, cheap and reproducible, but brittle when valid answers vary. Model-based graders, often called LLM-as-judge, handle open-ended output but need calibrating against human judgment. Human graders are the reference, and the slowest and most expensive. LangSmith separates offline evaluation, on curated datasets before release, from online evaluation of live traffic.

Two kinds of suite. Capability evals start with low pass rates and show progress. Regression evals should pass almost every time and catch breakage. For reliability, pass^k, where all k trials succeed, is a stricter target than pass@k, where at least one does.

Across organizations. When an agent works with another company’s agent, only the exchanged messages, task states and artifacts are visible from outside, so evaluations grade those and the resulting outcome.

Neighbouring terms. Tracing produces the transcripts that graders read.

Sources

  1. Anthropic: Demystifying evals for AI agents (9 January 2026) (accessed )
  2. Agent Development Kit documentation: Why evaluate agents (accessed )
  3. LangSmith documentation: Evaluation concepts (accessed )
  4. Anthropic: How we built our multi-agent research system (13 June 2025) (accessed )