arXiv:2609.30849v1 Announce Type: new Abstract: Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
A new arXiv paper (2609.30849v1) proposes an LLM-as-a-judge approach to measure perturbation strength uniformly across input and chain-of-thought (CoT) perturbations, enabling controlled-strength evaluation of self-consistency in LLM-generated explanations. The authors report their LLM-based perturbation strength measure outperforms embedding- and probability-based approaches, and that input perturbations generally affect LLMs more strongly than CoT perturbations. The paper concludes that judgments about a model's self-consistency are fair only within the same perturbation type.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.