{"slug": "knowing-the-form-not-the-function-automatically-auditing-answer-authority-in", "title": "Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks", "summary": "A study from arXiv (2608.02621v1) found that four large language models (LLMs) spontaneously cited legal authority in 238 Taiwan bar-examination items even when not prompted, but answer correctness and authority grounding often diverged: in criminal law, 24.0–42.4% of valid responses were answer-correct but missed the gold authority, while 15.2–21.7% were answer-incorrect but cited it. The authors propose joint answer–authority evaluation for statute-grounded legal benchmarks, arguing that answer-only scoring misclassifies naturally occurring authority misses as successes.", "body_md": "arXiv:2608.02621v1 Announce Type: new\nAbstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items. Because each item has a verified governing provision, we automatically audit answer correctness and authority grounding jointly. The two dimensions dissociate in both directions. In criminal law, 24.0--42.4\\% of valid responses were answer-correct but missed the gold authority, while 15.2--21.7\\% were answer-incorrect but cited it. A separate statutory-retrieval probe and a permissive citation-abstention intervention further show that answer and citation behavior can move separately at the output level. Because this mismatch arises without adversarial or inconsistency-inducing prompting, answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes. Because statutory authority is structurally extractable and externally verifiable, the failure can be measured automatically. A preliminary PRC civil-law extension also observes citation-unrequested authority marking, motivating a full cross-jurisdictional joint audit. We therefore propose joint answer--authority evaluation for statute-grounded legal benchmarks.", "url": "https://wpnews.pro/news/knowing-the-form-not-the-function-automatically-auditing-answer-authority-in", "canonical_source": "https://arxiv.org/abs/2608.02621", "published_at": "2026-08-05 04:00:00+00:00", "updated_at": "2026-08-05 04:03:27.283360+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-ethics"], "entities": ["arXiv", "Taiwan bar examination", "PRC civil law"], "alternates": {"html": "https://wpnews.pro/news/knowing-the-form-not-the-function-automatically-auditing-answer-authority-in", "markdown": "https://wpnews.pro/news/knowing-the-form-not-the-function-automatically-auditing-answer-authority-in.md", "text": "https://wpnews.pro/news/knowing-the-form-not-the-function-automatically-auditing-answer-authority-in.txt", "jsonld": "https://wpnews.pro/news/knowing-the-form-not-the-function-automatically-auditing-answer-authority-in.jsonld"}}