cd /news/large-language-models/enhancing-assessment-of-self-consist… · home › topics › large-language-models › article
[ARTICLE · art-140779] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength

A new arXiv paper (2609.30849v1) proposes an LLM-as-a-judge approach to measure perturbation strength uniformly across input and chain-of-thought (CoT) perturbations, enabling controlled-strength evaluation of self-consistency in LLM-generated explanations. The authors report their LLM-based perturbation strength measure outperforms embedding- and probability-based approaches, and that input perturbations generally affect LLMs more strongly than CoT perturbations. The paper concludes that judgments about a model's self-consistency are fair only within the same perturbation type.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.30849v1 Announce Type: new Abstract: Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/enhancing-assessment…] indexed:0 read:1min 2026-09-28 · —