cd /news/large-language-models/a-removal-based-approach-to-improve-… · home topics large-language-models article
[ARTICLE · art-121915] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

A Removal Based Approach to Improve LLM Faithfulness at Test-Time

Researchers introduced a test-time method that improves the faithfulness of large language model (LLM) explanations by removing input concepts not credited in the model's explanation and re-querying the model, targeting incompleteness in explanations. Across two datasets, multiple model families, and two independent faithfulness metrics, the approach outperformed standard prompting and prompting to encourage faithfulness, without modifying model parameters.

read1 min views1 publishedSep 7, 2026

arXiv:2609.04343v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model's decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model's answer. Existing approaches to improving LLM faithfulness include training-time methods, which require access to model weights and extensive computational resources, and test-time methods that largely focus on addressing unsoundness. We introduce a test-time approach that directly targets incompleteness. We remove from the input the concepts not credited in the model's explanation and re-query the model on the reduced input. This eliminates unmentioned influences while preserving the influence of mentioned concepts. Across two datasets, multiple model families, and two independent faithfulness metrics, our approach improves explanation faithfulness compared to both standard prompting and prompting to encourage faithfulness. Our method is model-agnostic and can be applied at inference time without modifying model parameters, providing a flexible mechanism for reducing hidden influences and improving the reliability and safety of LLM-assisted decision making.

── more in #large-language-models 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-removal-based-appr…] indexed:0 read:1min 2026-09-07 ·