cd /news/large-language-models/do-llms-make-more-mistakes-if-they-d… · home topics large-language-models article
[ARTICLE · art-125374] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Do LLMs Make More Mistakes If They Do Not Believe the Input Data?

A study posted to arXiv (2609.09363v1) found only a weak context-memory conflict in large language models, with counterfactual RDF triple inputs scoring just 0.05 lower on a 1-5 faithfulness scale than factual inputs when judged by Kimi K3. The researchers tested English, Czech, Slovak and Upper Sorbian text generation from factual, counterfactual and fictional triples containing local Czech and Slovak data, and reported that a suboptimal choice of LLM judge would overestimate the strength of the context-memory conflict. The finding runs contrary to the authors' expectations and bears on LLM usability in retrieval-augmented generation and data-to-text systems.

by read1 min views3 publishedSep 10, 2026

arXiv:2609.09363v1 Announce Type: new Abstract: Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input data based on local knowledge, only partially captured in models' parametric knowledge. We let the models generate text in English, Czech, Slovak and Upper Sorbian from factual (FA), counterfactual (CFA) and fictional (FI) RDF triples containing local Czech and Slovak data. Contrary to our expectations, we observe only a weak context-memory conflict on the human-annotated sample. For Kimi K3 as an LLM judge, which agrees well with human annotations on the sample, counterfactual inputs receive only slightly lower faithfulness scores than factual ones (-0.05 on a 1-5 scale). We also find that a suboptimal choice of LLM judge would lead to overestimating the strength of the context-memory conflict.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/do-llms-make-more-mi…] indexed:0 read:1min 2026-09-10 ·