cd /news/large-language-models/you-really-didn-t-get-that-benchmark… · home topics large-language-models article
[ARTICLE · art-121870] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments

A new benchmark for evaluating large language models' social pragmatic inference on Chinese online comments shows that the strongest model, GPT-4, achieves 81.42% leave-writer-out accuracy, while the mean across eight models is 68.70% compared to human accuracy of 90.8%. The benchmark, built from over 200,000 public Chinese social media records, includes 4,735 human-validated items pairing target comments with context and plausible misreadings, revealing that models often recognize irony or playfulness but misidentify the interactional move.

read1 min views1 publishedSep 7, 2026

arXiv:2609.04384v1 Announce Type: new Abstract: Chinese online comments often convey social meaning through indirect and playful language that is hard to interpret without context. Existing evaluations largely organize items around predefined phenomena or controlled pragmatic categories, leaving open whether models can distinguish plausible readings of what a naturally occurring comment is doing in a particular exchange. We introduce a benchmark for evaluating whether LLMs can recover such situated pragmatic meanings. From more than 200,000 public Chinese social media interaction records, we construct 4,735 human-validated diagnostic items, each pairing a target comment with reconstructed preceding context and plausible misreadings. We evaluate eight LLMs as both question writers and solvers in a cross-writer setting. The task is challenging: the strongest model achieves 81.42% leave-writer-out accuracy. Across all eight models, the mean leave-writer-out accuracy is 68.70% while human accuracy was 90.8%. Case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.

── more in #large-language-models 4 stories · sorted by recency
── more on @gpt-4 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/you-really-didn-t-ge…] indexed:0 read:1min 2026-09-07 ·