cd /news/large-language-models/what-is-good-extracting-and-testing-… · home topics large-language-models article
[ARTICLE · art-71416] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

What is Good? Extracting and Testing Implicit Theories of Literary Quality from LLM Reasoning Traces

A new study from arXiv (2607.20425v1) finds that reasoning-enabled LLMs evaluate literary quality with 79.3% accuracy across six quality tiers, prioritizing craft, depth, and distinctive voice over correctness. In systematic degradation tests, structure and voice loss caused the largest quality drops (2.78 and 2.34 points respectively), while vocabulary simplification had minimal impact (0.41 points), suggesting LLM judgments are holistic and more sensitive to structural than lexical features.

read1 min views1 publishedJul 24, 2026

arXiv:2607.20425v1 Announce Type: new Abstract: What makes writing "good" remains a persistent question in literary studies and computational linguistics. We present a two-study investigation of how reasoning-enabled LLMs evaluate literary quality. In Study 1, we construct a benchmark of 30 real texts spanning six quality tiers, from canonical literature to anonymous forum posts, and extract the model's implicit theory of quality from its reasoning traces. Across five DeepSeek replications, the model achieves 79.3% mean tier-classification accuracy. The traces reveal a consistent stated theory: the model values intentionality over correctness, prioritizing craft, depth, and distinctive voice. A familiarity experiment with style-matched but unrecognizable passages suggests that source recognition may inflate scores, although this is confounded by genuine quality differences between canonical originals and researcher-written pastiches. In Study 2, we probe this theory through systematic degradation of five canonical prose passages. We apply six manipulations - vocabulary simplification, rhythm flattening, imagery removal, voice genericization, structure simplification, and combined degradation - and reevaluate each version. Vocabulary simplification causes the smallest quality loss (0.41 +/- 0.46 points), far below structure (2.78) or voice (2.34) loss. Combined degradation is devastating (-5.64) but subadditive. An exploratory comparison with Qwen QwQ shows the same broad qualitative pattern. Together, these studies suggest that LLM judgments of writing quality are holistic, author-specific, and more sensitive to structural than lexical features, with implications for automated writing feedback and computational aesthetics.

── more in #large-language-models 4 stories · sorted by recency
── more on @deepseek 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/what-is-good-extract…] indexed:0 read:1min 2026-07-24 ·