{"slug": "when-do-llms-apply-the-wrong-law-diagnosing-llm-failures-in-temporal-legal", "title": "When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning", "summary": "A new arXiv study (2608.14610v1) finds that large language models (LLMs) exhibit a strong bias toward applying the most recently enacted law regardless of when the legally relevant facts occurred, and that this bias is not due to an inability to understand temporal scope or lack of knowledge about historical statutes. The researchers, who constructed a benchmark for temporal applicable-law determination, provide behavioral evidence that reinforcement-learning-shaped explicit reasoning reduces the diversity of reasoning paths, causing models to converge on current law, and that stronger general reasoning ability inversely correlates with performance on temporal legal reasoning.", "body_md": "arXiv:2608.14610v1 Announce Type: new\nAbstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination. However, whether large language models (LLMs) can reliably perform this task remains unexplored. In this paper, we construct a benchmark to evaluate LLMs on temporal applicable-law determination, and systematically investigate why they fail at temporal legal reasoning. Our experiments reveal four key findings. First, LLMs exhibit a strong bias toward applying the most recently enacted law, regardless of when the legally relevant facts occurred. Second, this bias does not stem from an inability to understand that laws have temporal scope, nor from a lack of knowledge about historical statutes. Third, we provide behavioral evidence that reinforcement-learning-shaped explicit reasoning may be a key mechanism: while improving general reasoning ability, it reduces the diversity of reasoning paths, causing models to converge on applying the current law. Fourth, this produces a counterintuitive inverse relationship: models with stronger general reasoning ability tend to perform worse on temporal legal reasoning. Our findings offer concrete guidance for future work on improving LLM performance in temporally grounded legal reasoning.", "url": "https://wpnews.pro/news/when-do-llms-apply-the-wrong-law-diagnosing-llm-failures-in-temporal-legal", "canonical_source": "https://arxiv.org/abs/2608.14610", "published_at": "2026-08-18 04:00:00+00:00", "updated_at": "2026-08-18 04:14:20.970476+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research"], "entities": ["arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-do-llms-apply-the-wrong-law-diagnosing-llm-failures-in-temporal-legal", "markdown": "https://wpnews.pro/news/when-do-llms-apply-the-wrong-law-diagnosing-llm-failures-in-temporal-legal.md", "text": "https://wpnews.pro/news/when-do-llms-apply-the-wrong-law-diagnosing-llm-failures-in-temporal-legal.txt", "jsonld": "https://wpnews.pro/news/when-do-llms-apply-the-wrong-law-diagnosing-llm-failures-in-temporal-legal.jsonld"}}