cd /news/large-language-models/same-text-different-numbers-the-dive… · home › topics › large-language-models › article
[ARTICLE · art-140864] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Same Text, Different Numbers: The Divergence of LLM-Based Measures

A study of thirteen LLM-based textual measures across seven LLMs from different providers found cross-model rank correlations averaging only 0.52, with transcript-level differences common across providers accounting for just 34% of total score variation. The research, posted as arXiv:2609.31013v1, scored earnings call transcripts of S&P 500 companies on constructs including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk, and found that model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. The authors conclude LLM-generated variables should be treated as model-contingent measurements and validated across providers.

by read1 min views1 publishedSep 28, 2026

arXiv:2609.31013v1 Announce Type: cross Abstract: Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/same-text-different-…] indexed:0 read:1min 2026-09-28 · —