arXiv:2609.22127v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to draft contractual language, yet conventional accuracy or preference-based evaluations are poorly matched to legal drafting. A clause may be fluent and stylistically polished while still omitting an essential carve-out, allocating risk in an unenforceable way, assuming an inapplicable jurisdiction, or exposing a party to regulatory liability. This paper presents a empirical study design and framework for evaluating LLM-generated contract clauses. The study evaluates four models - Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash, across 22 contract clause categories and 34 legally-motivated failure modes. We combine two evaluation frameworks: CLAUSE, which classifies prompts by legal function and failure target, and LENS-CRAFT, which scores outputs across nine legal-quality dimensions. Instead of averaging dimension scores, the study applies a Max Severity Principle so that a single legally decisive defect remains visible. The paper provides the evaluation protocol, taxonomy, analysis plan, and a results structure for reporting empirical findings. We argue that legal AI evaluation should move beyond aggregate accuracy toward clause-specific, failure-mode-driven, and risk-sensitive assessment.
Beyond Accuracy and Surface Fluency: Risk-Sensitive Evaluation of LLMs for Legal Clause Generation
A new arXiv paper proposes a risk-sensitive evaluation framework for large language models that draft legal contract clauses, arguing that conventional accuracy and preference-based benchmarks miss legally decisive defects. The study evaluates four models — Claude Haiku 4.5, Gemini 2.5 Flash Lite, GPT 5.4 Nano, and Qwen 3.5 Flash — across 22 contract clause categories and 34 legally-motivated failure modes, combining the CLAUSE prompt taxonomy with LENS-CRAFT's nine legal-quality dimensions. Rather than averaging dimension scores, the framework applies a Max Severity Principle so a single legally decisive defect remains visible, and the authors argue legal AI evaluation should move toward clause-specific, failure-mode-driven, risk-sensitive assessment.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.