Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark Harvey announced version 1.1 of LAB-AA, the legal agent benchmark run by independent evaluation firm Artificial Analysis, adding hallucination checks that score whether a model's legal outputs are explicitly supported by the source documents they cite. The update builds on the benchmark's private 120-task set launched around July 7, 2026, where top models post criterion pass rates of roughly 93-95% but all-pass rates often sit under 30% as of October 2026. Harvey separately reports its Harvey Assistant product has a 0.2% hallucination rate versus 0.7% to 1.9% for foundation models, its own vendor-reported figures. Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark The update scores AI models on whether their legal outputs are actually supported by the source documents they cite Legal AI just got a new kind of report card. Version 1.1 of LAB-AA, the legal agent benchmark tied to Harvey’s open-source framework, now includes hallucination checks. The idea is simple. If a model makes a claim, the benchmark wants to know whether the source documents back it up. What changed in LAB-AA v1.1 Harvey announced the v1.1 release of LAB-AA, which adds a layer of hallucination checking to how models are scored. Outputs are now judged on whether they have explicit support in the underlying source documents. LAB-AA is the version of the benchmark run by Artificial Analysis, the independent AI evaluation firm. It launched around July 7, 2026, using a private set of 120 tasks and a single-judge scoring system. Keeping the task set private serves a practical purpose. Models cannot be quietly tuned to ace questions they have never seen, which keeps the leaderboard honest. The new hallucination metric sits on top of the existing scoring approach. Previously, the focus was on whether a model’s answer satisfied expert-written rubric criteria. Now the benchmark also asks whether each claim is grounded in what the documents actually say. A benchmark built to be hard As of October 2026, the LAB-AA leaderboard shows top models posting criterion pass rates of roughly 93-95%. All-pass rates, which require a model to satisfy every criterion on a task, often sit under 30%. AI, tech, and the markets they move—in one daily briefing. Daily. Free. Join 34,000+ readers across crypto, finance, and policy. That gap is the whole point of all-pass scoring. Legal work tends to be binary in its consequences. A brief that nails nine arguments and botches the tenth can still lose the motion. Adding hallucination checks raises the bar further. A model could, in principle, hit rubric criteria while padding its answer with claims the documents do not support. The v1.1 update is designed to catch exactly that behavior. How LAB got here Harvey open-sourced its Legal Agent Benchmark, known as LAB, on May 6, 2026. The release included more than 1,200 tasks spread across 24-plus practice areas. Those tasks are graded against more than 75,000 rubric criteria written by legal experts. Artificial Analysis followed roughly two months later with LAB-AA, carving out its private 120-task subset and applying its own evaluation setup. On September 18, 2026, Harvey released lab-core v1.1.0, which introduced dual-judge evaluations alongside rubric improvements. Under that system, two AI models act as graders: Claude Sonnet 4.6 and GPT-5.5. Using two judges from different developers is meant to reduce the risk that one model’s quirks skew the scores. Note that lab-core v1.1.0 and LAB-AA v1.1 are separate releases. The first updated Harvey’s open-source core. The second updated the Artificial Analysis-run benchmark with hallucination scoring. Harvey’s own hallucination numbers Harvey has also published data on hallucination rates. According to the company, its Harvey Assistant product shows a hallucination rate of 0.2%. By Harvey’s figures, foundation models land between 0.7% and 1.9%. Those are the company’s own reported numbers, so they come with the usual caveat attached to any vendor measuring its own product. In a legal workflow that touches thousands of documents, the gap between 0.2% and 1.9% could translate into a meaningful difference in the number of errors a reviewing attorney has to catch. Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy https://cryptobriefing.com/editorial-policy/ .