cd /news/large-language-models/sword-wikidata-based-distortions-rev… · home topics large-language-models article
[ARTICLE · art-125372] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

SWORD: Wikidata-based Distortions Reveal Hidden Cross-Lingual Inconsistencies in LLM Factual Error Rejection

A new arXiv paper introduces SWORD (Systematic Wikidata-based Object-Relation Distortion), a benchmark that tests whether large language models consistently reject factually incorrect statements across eight widely spoken languages by perturbing Wikidata triples. The benchmark found cross-lingual performance gaps of up to 28 percentage points, a 49% relative reduction, on distorted statements in East Asian languages, and that models scored higher on semantically plausible distortions than on nonsensical random substitutions, indicating reliance on distributional familiarity rather than genuine factual verification. The authors state these asymmetric multilingual factual reasoning capabilities are obscured by conventional aggregate accuracy benchmarks.

by read1 min views2 publishedSep 10, 2026

arXiv:2609.09349v1 Announce Type: new Abstract: Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding. We introduce Systematic Wikidata-based Object-Relation Distortion (SWORD), a benchmark that evaluates whether models consistently reject factual errors across languages. SWORD generates syntactically well-formed but factually incorrect statements in eight widely spoken languages through controlled perturbations of Wikidata triples, ranging from random entity substitutions to semantically plausible property-based selections. Our distortion-based evaluation surfaces two critical insights that remain entirely obscured by conventional benchmarks. First, models counterintuitively achieve higher accuracy on semantically plausible distortions than on nonsensical random substitutions, suggesting reliance on distributional familiarity rather than genuine factual verification. Second, models exhibiting comparable baseline accuracy across languages show substantial performance degradation specifically on (East) Asian languages when presented with distorted statements, with cross-lingual performance gaps reaching up to 28 percentage points (49% relative reduction) in some models. These findings demonstrate that multilingual factual reasoning involves asymmetric capabilities that aggregate accuracy metrics systematically obscure.

── more in #large-language-models 4 stories · sorted by recency
── more on @sword 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/sword-wikidata-based…] indexed:0 read:1min 2026-09-10 ·