Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court A new arXiv paper (2609.26945v1) introduces a sentence-level benchmark for evaluating how well large language models classify interpretive canons, drawing on Larenz's conception of interpretation in the Savigny tradition and a dataset of German Federal Constitutional Court decisions annotated at the sentence level. Baseline evaluations of four LLMs from three model families under expert hand-written prompts yielded mean F1 scores between 70.4 and 79.2 across seven binary subtasks, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest. Prompts optimized with Genetic-Pareto (GEPA) did not systematically outperform the hand-written prompts under the tested configuration, which the authors say indicates the expert prompts provide a meaningful baseline. arXiv:2609.26945v1 Announce Type: new Abstract: Judicial reasoning remains challenging for large language models LLMs to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto GEPA . Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.