cd /news/ai-agents/ai-agent-evaluation-a-practical-guid… · home › topics › ai-agents › article
[ARTICLE · art-144625] src=pub.towardsai.net ↗ pub= topic=ai-agents verified=true sentiment=· neutral

AI Agent Evaluation: A Practical Guide with a Working Application

A local research assistant built on DeepAgents, LangGraph, ChromaDB and Llama 3.1 8B Instruct raised citation presence from 30.8% to 97.4% on answerable holdout turns after changes to reference handling, yet neither modified version replaced the baseline because combined evaluation surfaced workflow failures, unsupported references and answers withheld by validation. In one saved response a valid passage identifier accompanied an incorrect explanation of a technical term: the reference passed a code check and an LLM judge awarded full credit even though comparison with the passage showed the claim was unsupported. The guide argues agent evaluation must establish both whether intended behavior improved and whether the change introduced failures elsewhere, using deterministic checks, retrieval metrics, human review, LLM judging and regression testing.

by read21 min views1 publishedOct 3, 2026

The reliability of an AI agent depends on several connected processes, including the selection of tools, retrieval of relevant information, interpretation of evidence, and preparation of the final response. A change intended to improve one process may also affect the others. For a research assistant, requiring a source and page number may increase the frequency of citations without improving the support for the answer. The agent may cite a real passage that does not support its claim, omit a required verification step, or decline a question that the available evidence could answer. Evaluation therefore needs to establish both whether the intended behavior improved and whether the change introduced failures elsewhere in the application.

The local research assistant examined in this study provides a practical demonstration of that problem. Changes to reference handling increased citation presence from 30.8% to 97.4% on answerable holdout turns. However, neither modified version replaced the baseline because the combined evaluation identified workflow failures, unsupported references, and answers withheld by validation. In one saved response, a valid passage identifier accompanied an incorrect explanation of a technical term. The reference passed a code check, and an LLM judge awarded full credit, although comparison with the passage showed that the claim was unsupported. The following sections use these experiments to explain deterministic checks, retrieval metrics, human review, LLM judging, and regression testing. Particular attention is given to what each method measures and how its limitations affect the interpretation of agent performance.

The application answers questions about three papers: Attention Is All You Need, BERT, and the Llama 3 report. Requests range from a simple fact, such as the number of attention heads in the Transformer base model, to comparisons across papers. The Coordinator receives the question, delegates work to specialists, and prepares the final answer. The Retriever searches ChromaDB directly, without making an LLM call. The Fact Checker receives a proposed claim, retrieves evidence separately, and calls the model to assess that claim.

The papers are split within each PDF page using recursive character splitting, with a target size of 1,000 characters and 200 characters of overlap. The splitter tries to respect natural text boundaries, and each passage retains its source and PDF page number in ChromaDB. Llama 3.1 8B Instruct handles model calls through a local NVIDIA NIM container on an RTX 3090. DeepAgents and LangGraph organize the workflow, while Streamlit provides the interface.

For a supported technical question, the intended sequence is retrieval, claim checking, and answer preparation. A greeting should bypass the specialists, while an unsupported request should receive an explanation of the collection’s limits. These different paths provide observable requirements against which the application can be evaluated. Although the sequence is specific to the research assistant, the evaluation approach can be applied to other agent frameworks by examining the actions taken, the evidence available at each step, and the response delivered to the user. Agent evaluation begins with a description of the expected behavior under a defined set of conditions. For the question “How many attention heads does the Transformer base model use?”, the expected fact is eight. The research assistant is also required to retrieve supporting evidence, check a proposed claim, and cite the source. Together, these requirements form a behavior contract, which specifies what the application must do, what it may do, and what it should avoid. The contract extends beyond factual questions. A greeting requires a direct reply, a question about an unsupported paper requires an explanation of scope, and a follow-up may require information from an earlier turn. A request for BERT’s exact training electricity bill should receive an appropriate refusal because the presence of hardware information does not establish that cost.

These expectations are recorded with the input and starting conditions so that a failure can be interpreted against the intended task. For the attention-head question, a compact test case might contain the following fields:

Recording the evidence, workflow, and starting state makes it possible to distinguish failures that would be obscured by a question-and-answer pair alone. An answer could state eight but cite BERT, identify the correct paper but omit verification, or succeed because evidence remained from an earlier conversation. Reference labels also require a review status, since expectations drafted with AI assistance remain provisional until the intended human review has occurred. The same requirements apply when the proposed change concerns only citations or wording; the scope of the change does not remove the need to evaluate other required behaviors.

The project dataset contains 60 cases covering single-paper questions, comparisons, unavailable information, casual messages, follow-ups, and claim verification. Fifty-two cases exercise the full agent, while eight test the Fact Checker separately. Coverage comes from the behaviors represented, rather than the count alone. Repeating many easy factual questions would provide little evidence about refusal behavior, cross-paper comparisons, or conversation context.

A prompt revised repeatedly against the attention-head question may eventually produce the expected answer. That result demonstrates improvement on a known case, but provides limited evidence about performance on other questions. Separating development from evaluation helps address this limitation. A development set contains the cases used during inspection and revision, while a holdout set reserves other cases for assessing the chosen design after the prompts and grading rules have been fixed. The distinction is relevant to prompt changes as well as code changes because both can become tailored to the examples used during development.

The dataset was divided into 40 development cases and 20 holdout cases. Related paraphrases and conversation turns remained in the same set to reduce overlap between the conditions used for revision and those used for evaluation. Once a holdout failure informs a further revision, the affected example becomes part of development, and a later independent assessment requires fresh held-out cases. Maintaining that separation is necessary when interpreting whether an observed improvement extends beyond the questions used to produce it.

The study also has a limitation worth making explicit. An earlier diagnostic run stopped after a development case exposed a validation bug, but some holdout trials had already executed. The correction followed the development failure, and the final experiment restarted in fresh sessions, excluding the earlier results. Fresh sessions prevent conversation state from carrying over. They do not make previously exposed cases an untouched first evaluation, so the saved study record preserves that history.

Several requirements can be assessed directly from the execution record. Code can determine whether retrieval occurred, whether the Fact Checker followed retrieval, and whether a cited passage appeared in the evidence supplied to the agent. These are deterministic checks, meaning that a fixed rule applied to the same saved input produces the same result. The evaluation harness, which runs the cases and saves their outputs, applies these checks to recorded tool calls and evidence. The execution record is necessary because a statement in the final answer that a fact was verified does not establish that verification occurred.

The limitation of a structural check becomes apparent when an answer states “64 attention heads” while citing the correct page. The passage and page exist, and the required tools may have run, although the paper specifies eight heads. Semantic evaluation examines the meaning of the answer in relation to the evidence. A simple numerical requirement may still be checked by extracting and comparing the value, whereas a broader response may paraphrase several passages or omit a qualification that changes the conclusion. Such cases require interpretation by a human reviewer or a calibrated LLM evaluator. Structural and semantic results therefore need to remain distinguishable: reference presence establishes that a reference was supplied, while answer quality depends on whether the referenced evidence supports the claim.

Retrieval evaluation examines whether the information supplied to the agent is sufficient for the requested answer. A comparison of positional encoding in the original Transformer and Llama 3 papers, for example, requires evidence from both sources. If a search returns only the Transformer paper, source recall is one of two expected sources, or 50%. Returning five passages from that paper does not improve source recall because the second source remains absent. Finding both papers raises the score to 100%, although the retrieved passages may still omit the sections needed for the comparison.

More detailed checks are therefore needed to interpret evidence coverage. Reference-page recall measures whether known supporting pages were retrieved, with the qualification that useful evidence may also appear on pages outside the reference list. Passage relevance examines whether each returned excerpt contributes to answering the question. For example, a passage about attention may be topically related but provide no evidence about the number of attention heads. The distinction between source coverage and passage relevance is necessary when selecting metrics, since a retrieval method can perform well on the former while supplying insufficient evidence for the answer.

Relevance judgments support several metrics. Precision measures the useful share of returned passages: in an illustrative set of five results with two relevant passages, precision is 2/5, or 40%. Reciprocal rank measures where the first useful passage appears; a relevant result in second place receives 1/2. Averaging that score across questions gives mean reciprocal rank, or MRR. These calculations require appropriate relevance labels. Unreviewed passages should remain unjudged, and a report should define whether its cutoff counts results from one search or merged results from several searches.

Across 12 cross-paper cases, one vector search achieved 62.5% mean expected-source recall, while the application’s targeted searches reached 95.8%. The targeted method also returned an average of 9.4 passages instead of five. The experiment therefore compared two practical retrieval configurations without isolating the effects of query strategy and passage budget. A matched-budget comparison would help separate them. Broad passage relevance remained unreviewed, so the precision example above is a teaching calculation, not a measured project-wide result.

Workflow evaluation uses the tool trace to examine call order, arguments, results, and errors. In the research assistant, a supported technical request requires retrieval before the Fact Checker, and the checker must receive a declarative claim to assess. A greeting should bypass both specialists. These requirements describe the intended behavior of the application rather than a universal sequence for agents. A coding agent, for example, may allow several valid investigation paths while requiring tests before completion. Evaluation rules consequently need to reflect the application’s contract without rejecting alternative paths that satisfy the same requirements.

Conversation cases extend the evaluation to behavior that depends on previous turns. After “How many parameters does BERT Base have?”, the follow-up “And the larger version?” refers to BERT Large. Evaluating the follow-up in isolation changes the task and removes the context needed to interpret the request. The application’s scope guard checks the latest message for paper-related terms, allowing a follow-up to lose specialist access even when the earlier message has been stored correctly. The failure arises from the use of context during routing. Each conversation case therefore preserves its session across turns, while separate cases begin with empty memory. Examining these traces can reveal whether a citation-focused change also affects follow-ups, verification, or repeated tool use.

Human evaluation examines several dimensions of answer quality that may not agree for the same response. An answer stating “BERT Base has 110 million parameters” is factually correct; however, the answer is not supported by the available evidence if the agent received only a passage about masking policy. Correctness and grounding therefore require separate judgments. Completeness concerns coverage of the request, citation support concerns the references attached to claims, and appropriate abstention concerns the decision to answer or decline. A refusal may be justified when the collection lacks the requested fact but unjustified when the necessary evidence is available. Separating these dimensions makes the reason for a low score more informative than a single overall judgment.

Reviewers need both reference evidence, which establishes what a good answer should contain, and available evidence, which records what the agent actually received. In the citation experiments, available evidence means passages returned to the Coordinator by the Retriever, including retained passages from earlier turns. The Fact Checker’s private retrieval is recorded separately. Its generated verdict is not treated as an independent source document.

The rubric uses 2 for a pass, 1 for partial credit, and 0 for failure, with definitions specific to each dimension. For completeness, the scores correspond to covering all required facts, some, or none of the required answer. A central false claim fails correctness regardless of the fluency of the surrounding explanation. Citation support receives full credit when references support the substantive claims, partial credit for mixed support, and failure when references are absent or unrelated. An honest refusal without factual claims or misleading references can receive a not-applicable citation score. These definitions provide a common basis for assessment while preserving differences among factual errors, missing information, and unsupported claims.

The reviewer reads the question, unchanged answer, and evidence together, then records a reason for the judgment. One reviewed answer claimed that BERT uses only token embeddings. Completeness received partial credit for naming one component, while correctness failed because “only” excluded segment and position embeddings. Another answer used an unrelated passage ID alongside a supporting page number. The partial citation score becomes understandable only when the explanation records that conflict.

When several reviewers are available, independent judgments on a shared sample can reveal unclear instructions before disagreements are discussed. The completed calibration review in this study involved one human confirming visible AI-assisted draft scores, preferences, and reasons for 30 constructed examples and nine answer pairs. The review therefore represents human assessment with disclosed assistance; it does not provide blinded independent labels or a measure of agreement among human reviewers. Human assessment of the separate collection of 423 saved agent answers remains pending, which limits the conclusions that can be drawn about the quality of those responses.

An LLM judge applies a rubric to a saved answer. The grading input contains the question, candidate response, required facts, reference evidence, and evidence available to the agent. The prompt defines the scores and asks for reasons, while structured JSON allows code to reject missing fields, invalid score types, and unknown evidence IDs. Answers and source passages are treated as data, so instructions embedded inside them should not override the grading rules.

The offline judge has a different role from the application’s Fact Checker. The checker contributes to producing an answer; the judge grades the saved response afterward. Both used the same local Llama 3.1 8B model as the Coordinator, making the experiment feasible on one workstation while allowing shared misunderstandings. To examine that risk, the judge received 30 constructed examples containing correct responses and known defects, including false facts, missing citations, and unrelated references. Comparing the grades with reviewed examples is calibration.

Each example was graded three times. Twelve belonged to the holdout set, producing 36 ratings per applicable dimension. Agreement with the completed human review was:

Those are repeated judgments of 12 examples, not 36 independent examples. Six citation ratings were not applicable. The reviewed denominator also includes an appropriate refusal carrying an unrelated citation: declining the request was justified, but the attached reference remained misleading. The later human interpretation is preserved alongside the original provisional analysis, rather than silently replacing its labels.

One held-out answer correctly described the Transformer’s sine and cosine encodings but supplied no citation. The judge granted full citation credit in all three trials because it could see supporting evidence in its own input. It confused evidence available for grading with a reference supplied by the candidate. Across 21 holdout ratings expected to fail citation support, the judge awarded a full pass to 15. That is false acceptance, or approving an unacceptable result. False rejection is the reverse. Recording both directions explains risks that an average agreement rate can hide.

Pairwise evaluation compares two answers under a stated preference rule and allows the evaluator to select either answer or a tie. The judge evaluated nine reviewed pairs, with three repetitions and both display orders, producing 54 calls. Of those, 31 outputs were valid and all agreed with the human preference, while 23 failed validation. The agreement among valid judgments suggests that the judge could apply the preference rule to some comparisons, although the output failures substantially limited its operational reliability. Only 13 repeated comparisons had valid output in both orders, and none reversed preference; the small valid subset limits the evidence about sensitivity to answer order.

Swapping order tests whether placement changes a preference. Other probes can examine whether a judge rewards unnecessary length or obeys grading instructions embedded in an answer. Research on LLM-as-judge evaluation discusses position, verbosity, and self-enhancement biases. Our verbosity and instruction-injection probes each produced six invalid outputs. Those results establish an output-reliability problem, while leaving the intended preference behavior unresolved. Invalid output does not demonstrate resistance to bias or injection.

Repeated grading examines variability. All 30 rubric examples received identical score vectors across three trials in this experiment, including examples the judge scored incorrectly. Consistency therefore did not establish accuracy, and temperature zero should not be treated as a guarantee of identical future judgments. Calibration also provided little evidence about borderline cases because the reviewed collection contained only two partial-credit judgments.

Rubric wording introduced another limitation. The human review treats abstention as the decision to answer or decline, while grading factual errors separately. The frozen judge prompt also penalizes invention under abstention, so a false answer to an answerable question can receive different treatment under the two interpretations. Abstention agreement is excluded from the headline results until a new rubric version resolves the mismatch. The combined findings support using the judge to assist inspection, but not allowing its citation scores to approve an agent change automatically.

A regression occurs when a change damages behavior that previously met the application’s requirements. The citation experiment examined that risk by comparing three versions of the assistant. A, the baseline, allows the model to write citations in prose. B requires the model to link claims to stable passage IDs, after which code renders the source and page information. C adds reference validation and at most one repair attempt, withholding the answer if validation still fails. The progression introduces greater control over reference handling, but also creates additional work and the possibility of blocking answers that the evidence could support.

Each version received the same cases, retrieval settings, scope rules, model, and generation settings. The resulting paired experiment compares outcomes for the same case before summarizing across cases, making improvements and regressions attributable to particular tested conditions. Each case ran three times per version, and related trials remained grouped when estimating uncertainty. The 17 holdout agent cases, including one two-turn conversation, produced 54 technical turns per version. Citation presence was assessed on the 39 answerable turns so that appropriate refusals to unavailable requests were not penalized for lacking references.

Recognized citations appeared in 12 of 39 answerable turns for A, 32 for B, and 38 for C. Mean response time across all 54 technical turns increased from 6.36 seconds to 7.21 and 7.29 seconds, respectively. The results demonstrate that stronger control over references increased citation presence at an additional time cost. However, the citation count alone cannot establish whether the evidence supported the answer. Examination of the saved responses was therefore necessary to determine whether the structural improvement corresponded to a semantic improvement.

In the first trial comparing positional encoding, A and B both expanded RoPE incorrectly as “Reformer-style Positional Embeddings.” RoPE means Rotary Position Embedding, as defined in the RoFormer paper. B attached a valid passage ID from PDF page 7 of the Llama 3 report. The cited row identifies “RoPE (θ = 500,000),” supporting the use of RoPE and the parameter value but supplying no expansion of the acronym. The reference check passed because the identifier belonged to the available evidence set. Membership alone could not establish the meaning of the claim.

Version C withheld its answer after a repair attempt failed structural validation. The block establishes failure to satisfy the reference rules, but does not demonstrate that the validator identified the incorrect expansion. The offline judge nevertheless granted full credit on every dimension to all three responses, including C’s message, “Unable to produce an answer with valid evidence references.” The grading explanation credited facts from the reference material as though they appeared in the candidate answer. Comparison of the response, cited passage, and grading reason thus identified separate failures in the agent and the evaluator, neither of which was apparent from the citation-presence score.

Across development and holdout, C attempted 25 repairs and withheld 12 answers. Some affected cases already had poor baseline answers, so all blocks cannot be called new semantic regressions. Neither candidate was promoted: B introduced a development routing failure and referenced passages outside the Coordinator’s supplied evidence in two other trials; C also introduced a routing failure and triggered the blocking rule. B had no new failures under the holdout code checks, but known development problems still mattered. Keeping A as the comparison baseline did not certify its reliability.

An improvement in answer quality may justify additional response time or resource use, depending on the requirements of the application. Assessing that tradeoff requires measurements of all serving calls, including the Fact Checker’s internal call and any repair attempts, together with input and output tokens, elapsed time, and errors. Missing usage data must remain unknown rather than being recorded as zero. Greetings also need to be separated from complete research workflows, since combining fast direct replies with technical tasks can understate the time required for a full retrieval-and-checking sequence.

Offline evaluation has a separate cost. Calibration used 144 model calls, while grading 423 saved agent answers required another 423 calls, more than a million input tokens, and 31.76 minutes of cumulative model-call time. Forty-six grading responses failed validation. Those calls did not contribute to the user’s wait for an answer, but they affect the practicality of repeating the evaluation. On a local workstation, token counts describe workload; a monetary estimate also needs assumptions or measurements for hardware, electricity, and utilization.

Release criteria provide a basis for interpreting the tradeoff and should be defined before the final comparison is inspected. The project rules block new critical failures such as invented references, broken routing, tool loops, or execution errors, while latency increases above 20% and token increases above 25% trigger review. These thresholds reflect choices for the research assistant rather than general standards for agent deployment. Code checks contribute evidence to the release decision, but semantic review remains necessary. In particular, the completed human calibration review establishes judgments for the constructed examples and does not substitute for reviewing the separate collection of agent answers.

Interpretation of an evaluation result depends on the ability to inspect the observations underlying the reported score. The study preserves exact answers, retrieved passages, tool calls, grades, and reasons alongside configuration snapshots. Each result can therefore be traced to its case, repeated trial, agent version, and evaluation record. Later human judgments remain separate from the original labels so that changes in interpretation can be distinguished from changes in agent behavior. Preserving those records also allows an evaluator failure to be investigated without repeating the original agent run.

The companion materials on GitHub provide three sample cases, the human-review rubric, the tested judge prompt, and the saved A/B/C responses for the position-encoding example. From the repository root, run node examples/agent-evaluation/check-example.cjs. The example checks tool order and citation identifiers using Node.js, without a GPU, model server, or API key. Its README explains why a structural pass does not establish semantic support. The package demonstrates saved-output evaluation rather than reproducing the complete study, and its frozen rubric and prompt retain the limitations discussed above.

All measured results here concern offline evaluation. After release, online monitoring would examine real requests through representative traces, errors, latency, unexpected refusals, and human feedback. Investigated failures can become development cases and later regression tests. Product outcomes such as acceptance, completion, abandonment, adoption, and time saved add evidence about practical value, but each needs a clear definition. A user leaving the chat can indicate either success or frustration. These outcomes were not measured in our local study; Anthropic’s agent-evaluation guide discusses how evaluation extends to real operating conditions.

The results of this study demonstrate that improvements in a limited measure of agent behavior can coexist with failures elsewhere in the application. Reference validation increased citation presence, but the saved responses showed that valid identifiers could accompany unsupported claims and that an LLM judge could repeat the same error. Deterministic checks, retrieval metrics, human assessment, and judge calibration supplied different parts of the evidence needed to interpret those outcomes. Repeated paired experiments then connected the observed improvements and regressions to the proposed changes and their resource requirements. For practical agent development, the implication is that an improvement should be assessed against the combined behavior contract of the application, with the underlying responses and traces available for review before a release decision is made.

AI Agent Evaluation: A Practical Guide with a Working Application was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.

── more in #ai-agents 4 stories · sorted by recency
── more on @deepagents 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/ai-agent-evaluation-…] indexed:0 read:21min 2026-10-03 · —