{"slug": "beyond-accuracy-and-surface-fluency-risk-sensitive-evaluation-of-llms-for-legal", "title": "Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation", "summary": "A new arXiv paper proposes a risk-sensitive evaluation framework for large language models that draft legal contract clauses, arguing that conventional accuracy and preference-based benchmarks miss legally decisive defects. The study evaluates four models — Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash — across 22 contract clause categories and 34 legally-motivated failure modes, combining the CLAUSE prompt taxonomy with LENS-CRAFT's nine legal-quality dimensions. Rather than averaging dimension scores, the framework applies a Max Severity Principle so a single legally decisive defect remains visible, and the authors argue legal AI evaluation should move toward clause-specific, failure-mode-driven, risk-sensitive assessment.", "body_md": "arXiv:2609.22127v1 Announce Type: new \nAbstract: Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A clause may be fluent and stylistically polished while still omitting an essential carve-out, allocating risk in an unenforceable way, assuming an inapplicable jurisdiction, or exposing a party to regulatory liability. This paper presents a empirical study design and framework for evaluating LLM-generated contract clauses. The study evaluates four models - Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash, across 22 contract clause categories and 34 legally-motivated failure modes. We combine two evaluation frameworks: CLAUSE, which classifies prompts by legal function and failure target, and LENS-CRAFT, which scores outputs across nine legal-quality dimensions. Instead of averaging dimension scores, the study applies a Max Severity Principle so that a single legally decisive defect remains visible. The paper provides the evaluation protocol, taxonomy, analysis plan, and a results structure for reporting empirical findings. We argue that legal AI evaluation should move beyond aggregate accuracy toward clause-specific, failure-mode-driven, and risk-sensitive assessment.", "url": "https://wpnews.pro/news/beyond-accuracy-and-surface-fluency-risk-sensitive-evaluation-of-llms-for-legal", "canonical_source": "https://arxiv.org/abs/2609.22127", "published_at": "2026-09-22 04:00:00+00:00", "updated_at": "2026-09-22 04:25:32.918212+00:00", "lang": "en", "topics": ["large-language-models", "ai-safety", "ai-ethics", "ai-research"], "entities": ["arXiv", "Claude Haiku 4.5", "Gemini 2.5 Flash Lite", "GPT 5.4 Nano", "Qwen 3.5 Flash", "CLAUSE", "LENS-CRAFT"], "alternates": {"html": "https://wpnews.pro/news/beyond-accuracy-and-surface-fluency-risk-sensitive-evaluation-of-llms-for-legal", "markdown": "https://wpnews.pro/news/beyond-accuracy-and-surface-fluency-risk-sensitive-evaluation-of-llms-for-legal.md", "text": "https://wpnews.pro/news/beyond-accuracy-and-surface-fluency-risk-sensitive-evaluation-of-llms-for-legal.txt", "jsonld": "https://wpnews.pro/news/beyond-accuracy-and-surface-fluency-risk-sensitive-evaluation-of-llms-for-legal.jsonld"}}