cd /news/large-language-models/polite-but-misaligned-evaluating-llm… · home › topics › large-language-models › article
[ARTICLE · art-139418] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms

A study of seven large language models found that the models agree with each other on politeness judgments more than they agree with human raters, according to the arXiv paper 2609.29001v1. The paper reports that model-human alignment is associated with explicit linguistic cues, while some rapport-building strategies appear more often in misaligned cases, and that in the three-way categorical task models systematically overproduce Neutral labels and underpredict Impolite labels, a pattern that persisted against expert consensus on a diagnostic subset. The authors conclude that pragmatic evaluations should examine directional patterns of model-human disagreement rather than aggregate agreement metrics alone.

by read1 min views2 publishedSep 25, 2026

arXiv:2609.29001v1 Announce Type: new Abstract: Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that model--human alignment is associated with explicit linguistic cues, while some rapport-building strategies occur more frequently in misaligned cases. In the categorical task, model predictions exhibit systematic neutral compression, characterized by the overproduction of Neutral labels and the underprediction of Impolite labels. This pattern persists when expert consensus is used as the reference on a diagnostic subset. Our findings highlight the need for pragmatic evaluations that go beyond aggregate agreement metrics by examining directional patterns of model--human disagreement across different human references.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/polite-but-misaligne…] indexed:0 read:1min 2026-09-25 · —