cd /news/large-language-models/count-evidence-not-sentences-tempere… · home topics large-language-models article
[ARTICLE · art-138833] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement

Researchers proposed Tempered Evidence Fusion (TEF), a training-free decision-fusion rule that weights each sentence's log-odds by its normalized information gain, for measuring value orientations in long social media posts with large language models. On the new Multi-event Insight Network Dimensions (MIND) benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions, TEF outperformed the strongest of Direct, Majority Vote, and Soft Vote baselines by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. The MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.27165v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.

── more in #large-language-models 4 stories · sorted by recency
── more on @tempered evidence fusion 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/count-evidence-not-s…] indexed:0 read:1min 2026-09-24 ·