{"slug": "intlawner-a-named-entity-recognition-dataset-and-benchmark-in-international-law", "title": "IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law", "summary": "Researchers introduced IntLawNER, a named entity recognition dataset and benchmark for international law covering 2,987 gold-annotated sentences and 8,094 entity spans drawn from International Court of Justice decisions, UN Security Council resolutions, and European Court of Human Rights judgments. The dataset was built with a hybrid algorithmic-agentic pipeline that reduced 468,000 source sentences to a compact annotation set, with 89.6% of gold spans accepted unchanged from the silver layer, and the benchmark found zero-shot GLiNER collapses to 0.243 micro-F1 on institution-dependent entity types while Claude Opus 4.6 reached the best score of 0.873 micro-F1 with few-shot prompting. The authors report that human-machine agreement metrics can mislead in domain-specific NER, since Cohen's kappa of 0.964 on boundary-matched spans masks a macro-F1 of 0.753 once missing entities, boundary errors, and label corrections are included.", "body_md": "arXiv:2609.22529v1 Announce Type: new \nAbstract: International law provides the normative framework through which states coordinate action, regulate armed conflict, and protect human rights, yet its texts remain without token-level named entity recognition (NER) resources. We introduce IntLawNER, a NER dataset and benchmark for codified sources of international law, covering 2,987 gold-annotated sentences and 8,094 entity spans from International Court of Justice (ICJ) decisions, UN Security Council resolutions, and European Court of Human Rights (ECtHR) judgments, annotated with seven institution-specific entity types. We construct IntLawNER with a cost-effective hybrid algorithmic-agentic pipeline that reduces 468k source sentences to a compact annotation set through candidate retrieval, LLM-based vetting, and human review, with 89.6% of gold spans accepted unchanged from the silver layer. However, the silver-to-gold analysis reveals that human-machine aggregate agreement metrics can be misleading in domain-specific NER: Cohen's kappa=0.964 on boundary-matched spans masks a macro-F1 of 0.753 when missing entities, boundary errors, and label corrections are included. The benchmark shows that zero-shot span-based GLiNER collapses on entity types dependent on institutional function rather than surface form (0.243 micro-F1), while fine-tuned transformers struggle on rare labels. Carefully selected few-shot examples that demonstrate label contrasts improve every LLM over zero-shot prompting, with Claude Opus 4.6 reaching the best score of 0.873 micro-F1. We release IntLawNER as a benchmark and reusable resource for extracting references in international legal texts.", "url": "https://wpnews.pro/news/intlawner-a-named-entity-recognition-dataset-and-benchmark-in-international-law", "canonical_source": "https://arxiv.org/abs/2609.22529", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:26:38.670578+00:00", "lang": "en", "topics": ["natural-language-processing", "ai-research", "large-language-models", "ai-tools"], "entities": ["IntLawNER", "International Court of Justice", "UN Security Council", "European Court of Human Rights", "GLiNER", "Claude Opus 4.6"], "alternates": {"html": "https://wpnews.pro/news/intlawner-a-named-entity-recognition-dataset-and-benchmark-in-international-law", "markdown": "https://wpnews.pro/news/intlawner-a-named-entity-recognition-dataset-and-benchmark-in-international-law.md", "text": "https://wpnews.pro/news/intlawner-a-named-entity-recognition-dataset-and-benchmark-in-international-law.txt", "jsonld": "https://wpnews.pro/news/intlawner-a-named-entity-recognition-dataset-and-benchmark-in-international-law.jsonld"}}