The machine never raises its voice In a blind comparison, three AI models preferred machine-generated literary passages over works by famous authors more than 90% of the time, with DeepSeek V4 Flash choosing machine text in 94% of decided pairs, Claude Haiku 4.5 in 98%, and GPT-5.6 Luna in 96%. The experiment, run by a developer building a prose editing tool, used Claude Opus to generate eight literary paragraphs and eight short poems, then asked the models to judge them against passages from Austen, Dickens, Woolf, Kipling, Shakespeare, Milton, and others across 60 pairs, with order reversal to remove position bias. A larger model, Claude Opus 5, also favored machine text in 88% of decided pairs, with only a Shakespeare sonnet and works by Wordsworth, Dickens, Kipling, and Blake winning some human judgments. The machine never raises its voice On this page I have been building a prose editing tool, so I got to spend a lot of time watching a model rewrite sentences. After a while, I started to notice that it rewrote all sentences in a similar direction. Why did it keep doing that? Is there at the risk of humanizing the model a personal preference and some sense of an aesthetic there that its writing is drawn to? To answer this, I tried to figure out what an LLM considers good writing. The Experiment The idea is simple. Ask an LLM to compare text written by famous authors in English against text written by a machine. I used Claude Opus to write eight literary paragraphs and eight short poems. I also collected passages from famous real novelists and poets — Austen, Dickens, Woolf, Kipling, Shakespeare, Milton and more. Then I asked three different models GPT-5.6 Luna, DeepSeek V4 Flash and Claude Haiku 4.5 the same question: The passages below are literary writing: literary fiction and poetry. Which of these two passages is the better piece of literary writing? Answer with X or Y on the first line, then one short sentence saying why. Do not explain anything else. X