cd /news/large-language-models/knowing-the-form-not-the-function-au… · home topics large-language-models article
[ARTICLE · art-87110] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

A study from arXiv (2608.02621v1) found that four large language models (LLMs) spontaneously cited legal authority in 238 Taiwan bar-examination items even when not prompted, but answer correctness and authority grounding often diverged: in criminal law, 24.0–42.4% of valid responses were answer-correct but missed the gold authority, while 15.2–21.7% were answer-incorrect but cited it. The authors propose joint answer–authority evaluation for statute-grounded legal benchmarks, arguing that answer-only scoring misclassifies naturally occurring authority misses as successes.

read1 min views1 publishedAug 5, 2026

arXiv:2608.02621v1 Announce Type: new Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy for authority grounding. Under ordinary reasoning prompts that did not request statutory citations, four LLMs spontaneously produced authority markers across 238 Taiwan bar-examination items. Because each item has a verified governing provision, we automatically audit answer correctness and authority grounding jointly. The two dimensions dissociate in both directions. In criminal law, 24.0--42.4% of valid responses were answer-correct but missed the gold authority, while 15.2--21.7% were answer-incorrect but cited it. A separate statutory-retrieval probe and a permissive citation-abstention intervention further show that answer and citation behavior can move separately at the output level. Because this mismatch arises without adversarial or inconsistency-inducing prompting, answer-only scoring treats naturally occurring gold-authority misses as complete benchmark successes. Because statutory authority is structurally extractable and externally verifiable, the failure can be measured automatically. A preliminary PRC civil-law extension also observes citation-unrequested authority marking, motivating a full cross-jurisdictional joint audit. We therefore propose joint answer--authority evaluation for statute-grounded legal benchmarks.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/knowing-the-form-not…] indexed:0 read:1min 2026-08-05 ·