{"slug": "legal-llm-hallucination-should-be-evaluated-as-failure-of-legal-warrant", "title": "Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant", "summary": "A position paper posted to arXiv (2609.17546v1) argues that hallucinations in legal LLMs should be evaluated as a failure of legal warrant rather than as factual inaccuracy or citation failure. The authors define claim-authority warrant as the context-sensitive relation between a consequential legal claim and authority that exists, applies to the relevant jurisdiction, is current for the date of analysis, has the legal status represented by the system, and supports the proposition asserted, and predict that warrant metrics will reveal material failures missed by answer accuracy, citation existence, generic attribution, LegalHalBench-style statute relevance, and CitaLaw-style sentence-citation alignment. The paper specifies benchmark records, claim boundaries, support labels, mixed response-policy scoring, risk weights, annotation reliability reporting, and jurisdiction-specific authority ontologies, and includes a side-by-side comparison item and a small reproducible pilot over public-rule tests.", "body_md": "arXiv:2609.17546v1 Announce Type: new \nAbstract: In this position paper, we argue that legal LLMs' hallucinations should be evaluated as a failure of legal warrant rather than as factual inaccuracy or citation failure. We define claim-authority warrant as the context-sensitive relation between a consequential legal claim and authority that exists, applies to the relevant jurisdiction, is current for the date of analysis, has the legal status represented by the system, and supports the proposition asserted. Warranted legal generation is the broader system behavior that answers, narrows, asks, warns, corrects a false premise, or abstains according to that relation. The falsifiable prediction is that warrant metrics reveal material failures that answer accuracy, citation existence, generic attribution, LegalHalBench-style statute relevance, and CitaLaw-style sentence-citation alignment can miss. We sharpen this claim with a side-by-side comparison item and a small, reproducible pilot over public-rule tests. We then specify benchmark records, claim boundaries, support labels, mixed response-policy scoring, risk weights, annotation reliability reporting, and jurisdiction-specific authority ontologies. The result is a concrete research agenda for evaluating legal AI systems by whether their consequential claims are licensed by law.", "url": "https://wpnews.pro/news/legal-llm-hallucination-should-be-evaluated-as-failure-of-legal-warrant", "canonical_source": "https://arxiv.org/abs/2609.17546", "published_at": "2026-09-17 04:00:00+00:00", "updated_at": "2026-09-17 04:25:12.175287+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-ethics", "ai-research"], "entities": ["arXiv", "LegalHalBench", "CitaLaw"], "alternates": {"html": "https://wpnews.pro/news/legal-llm-hallucination-should-be-evaluated-as-failure-of-legal-warrant", "markdown": "https://wpnews.pro/news/legal-llm-hallucination-should-be-evaluated-as-failure-of-legal-warrant.md", "text": "https://wpnews.pro/news/legal-llm-hallucination-should-be-evaluated-as-failure-of-legal-warrant.txt", "jsonld": "https://wpnews.pro/news/legal-llm-hallucination-should-be-evaluated-as-failure-of-legal-warrant.jsonld"}}