Writing Evals for Agentic Systems as a 4 Step Loop A practitioner's guide published October 4, 2026 outlines a four-step loop for writing evaluations of agentic AI systems: freeze and version every agent run's prompts, files, tools and MCP responses; have a subject matter expert manually grade each step as correct, incorrect or extra with reasoning; compile roughly 20 SME-evaluated runs into a dataset; and convert recurring failures into a pass/fail rubric. The author recommends running deterministic code checks first — such as whether the database ended in the right state, the right tool was called, or output parsed against the schema — and reserving an LLM judge only for what code cannot check, noting that even at 0 temperature LLMs are not fully deterministic because inference details like batch size shift results. Writing Evals for Agentic Systems as a 4 Step Loop October 4, 2026 1043 words Evaluation of any system requires an input, a system logic and an output. The output is then compared to the desired output. This process forms the base of any form of evaluation in computers. In deterministic systems, say an arithmetic operation or querying a database, the output is deterministic i.e. it falls in a fixed domain. In non-deterministic systems like an LLM, the output can lie in an open domain. The distribution of outputs will be biased towards the desired output but it will never be 100%. Even at 0 temperature, LLMs are not fully deterministic in practice since inference level details like batch size can shift results slightly. Deterministic evaluations can usually be automated once created. These are unit tests, integration tests etc. They evaluate outputs against a defined set of expected behaviours or values. Non-deterministic systems cannot be evaluated purely by comparing their outputs to one fixed expected output, as their outputs can lie in an open domain. An LLM can output the same sentence in 100s of different ways, structures and details. Correctness of the system in this case relies on the criteria defined by the person in charge. Someone may like a brief response while someone wants every nitty detail with citations, someone may like to get information in bullet points while others may prefer paragraphs. Also not everything about an agent is open ended. Even when the wording varies, a lot of what it does can still be checked with plain code like did the database end up in the right state, was the right tool called, did the output parse against the schema. I always put these code checks first and only bring in a judge for what code can't check. Evaluating AI, more specifically agentic systems is tricky as they are non-deterministic systems and also very open-ended. The ideal output can be subjective and no universal guideline exists that determines what an ideal output looks like. I follow a 4 step process to write evaluations for my AI or Agentic systems: 1. Every operation performed by the agent needs to be stored with a frozen snapshot and should be versioned A trace of agent’s order of operation can look like All this needs to be stored as a frozen snapshot when initializing the agent run. All the system prompts should be versioned, the files, tools etc. available at the time of initialization need to be stored or otherwise be reproducible. A common mistake I did earlier was that I did not store the files and tools as frozen versions so if that file or tool is modified or deleted in future, then the old agent run can’t be replayed and thus makes it stale. The same goes for external tools and MCP calls, record their responses so you can mock them on replay. This should be done for all runs. 2. Evaluate the run manually This is the part where people stop building their AI evals as this is the most boring and manual part. A SME Subject matter expert needs to sit and evaluate the agent run we stored earlier. This step’s scope varies a lot. Ideally each step, each file read, tool output etc. needs to be evaluated when the intermediate behaviour also matters, to identify at what point the system loosens. Every relevant step should be marked as correct / incorrect / extra etc A fixed set of feedbacks can be designed and a comment. The comment should include reasoning on why a certain feedback is given. 3. Create a dataset of SME evaluated agent runs A dataset should be compiled which contains the goal, agent run and SME’s evaluation with the annotations. A dataset of around 20 is a good start which should grow with more diverse runs and new failure cases. Once you have enough graded runs, go through the SME comments and note the failures that keep repeating. Then those repeating failures should be turned into a rubric scoring guide of what counts as a good or a bad run i.e. a short list of pass/fail checks for what a good run looks like. This should be made after SME's grading not before, since your idea of good changes once you see real runs. 4. Make an LLM-as-a-Judge based on that dataset This step is where we are trying to reduce the SME’s work with another LLM. This LLM gets the rubric and a few labeled examples from the dataset in its prompt and will be used to evaluate the future runs. LLM-as-a-Judge should ideally be a more capable reasoning model, and its evaluations should be validated against SME or other reliable evaluations before being trusted at scale. LLM as a Judge can be used in 2 ways: Let LLM as a Judge evaluate a run on the fly i.e. while an agent runs, the judge keeps evaluating and potentially feeding feedback back into it. This can help catch undesired or less accurate outputs earlier, but increases the time taken to produce the output as LLM-as-a-judge adds another model call to the process. LLM-as-a-judge runs after a run. This can be a weekly run or an hourly cron job. A set of runs are picked and ran against the judge and the evaluation criteria. New failures and learnings can then be added back into the dataset. This is the general choice for most cases. LLM as a judge once created should not be thought of as a one time task, It needs to be monitored regularly and should be improved. Evaluator drift is a somewhat common tendency where the model used for the judge starts to diverge from SME preferences and this small drift can stack up over time. These learnings are then used to improve prompts, tool structure, addition of new guardrails and tools and to determine which LLM is best fit for the purpose. Everytime you change a prompt, add or update or delete a tool, change model, add or remove an MCP or change the harness, you should run the dataset created to check for regressions. This is also a good time to recheck the judge against the SME labels, especially if the new failures look different from the old ones.