cd /news/large-language-models/same-question-different-answers-eval… · home topics large-language-models article
[ARTICLE · art-76380] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

A new study from arXiv (2607.22554v1) finds that large language models frequently change their answers when the same question is rephrased, with mismatch rates exceeding 23% across four benchmarks and 13 models. The researchers show that while overall accuracy shifts modestly, instance-level behavior is unstable, and a simple self-paraphrasing strategy can recover latent knowledge and improve inference-time performance.

read1 min views1 publishedJul 28, 2026

arXiv:2607.22554v1 Announce Type: new Abstract: Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/same-question-differ…] indexed:0 read:1min 2026-07-28 ·