cd /news/large-language-models/beyond-accuracy-and-surface-fluency-… · home topics large-language-models article
[ARTICLE · art-136627] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation

A new arXiv paper proposes a risk-sensitive evaluation framework for large language models that draft legal contract clauses, arguing that conventional accuracy and preference-based benchmarks miss legally decisive defects. The study evaluates four models — Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash — across 22 contract clause categories and 34 legally-motivated failure modes, combining the CLAUSE prompt taxonomy with LENS-CRAFT's nine legal-quality dimensions. Rather than averaging dimension scores, the framework applies a Max Severity Principle so a single legally decisive defect remains visible, and the authors argue legal AI evaluation should move toward clause-specific, failure-mode-driven, risk-sensitive assessment.

by read1 min views1 publishedSep 22, 2026

arXiv:2609.22127v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A clause may be fluent and stylistically polished while still omitting an essential carve-out, allocating risk in an unenforceable way, assuming an inapplicable jurisdiction, or exposing a party to regulatory liability. This paper presents a empirical study design and framework for evaluating LLM-generated contract clauses. The study evaluates four models - Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash, across 22 contract clause categories and 34 legally-motivated failure modes. We combine two evaluation frameworks: CLAUSE, which classifies prompts by legal function and failure target, and LENS-CRAFT, which scores outputs across nine legal-quality dimensions. Instead of averaging dimension scores, the study applies a Max Severity Principle so that a single legally decisive defect remains visible. The paper provides the evaluation protocol, taxonomy, analysis plan, and a results structure for reporting empirical findings. We argue that legal AI evaluation should move beyond aggregate accuracy toward clause-specific, failure-mode-driven, and risk-sensitive assessment.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/beyond-accuracy-and-…] indexed:0 read:1min 2026-09-22 ·