{"slug": "enhancing-assessment-of-self-consistency-in-llm-explanations-using-perturbation", "title": "Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength", "summary": "A new arXiv paper (2609.30849v1) proposes an LLM-as-a-judge approach to measure perturbation strength uniformly across input and chain-of-thought (CoT) perturbations, enabling controlled-strength evaluation of self-consistency in LLM-generated explanations. The authors report their LLM-based perturbation strength measure outperforms embedding- and probability-based approaches, and that input perturbations generally affect LLMs more strongly than CoT perturbations. The paper concludes that judgments about a model's self-consistency are fair only within the same perturbation type.", "body_md": "arXiv:2609.30849v1 Announce Type: new \nAbstract: Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.", "url": "https://wpnews.pro/news/enhancing-assessment-of-self-consistency-in-llm-explanations-using-perturbation", "canonical_source": "https://www.machinebrief.com/news/enhancing-assessment-of-self-consistency-in-llm-explanations-mhhi", "published_at": "2026-09-28 04:00:00+00:00", "updated_at": "2026-09-28 04:48:13.237103+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "natural-language-processing", "artificial-intelligence"], "entities": ["arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/enhancing-assessment-of-self-consistency-in-llm-explanations-using-perturbation", "markdown": "https://wpnews.pro/news/enhancing-assessment-of-self-consistency-in-llm-explanations-using-perturbation.md", "text": "https://wpnews.pro/news/enhancing-assessment-of-self-consistency-in-llm-explanations-using-perturbation.txt", "jsonld": "https://wpnews.pro/news/enhancing-assessment-of-self-consistency-in-llm-explanations-using-perturbation.jsonld"}}