{"slug": "why-is-ai-bad-at-writing", "title": "Why Is AI Bad at Writing?", "summary": "AI-generated writing fails because it lacks the rhythm and flow of human composition, not because of surface tells like the em dash, according to a Weighty Thoughts essay by a writer who has used Claude for drafting but found the output required near-total rewriting. The author notes that AI-detection tools such as Pangram will need continuous retraining because model habits keep changing, and that different models do not all write the same. The essay argues the dislike of AI writing is often driven by the mere fact that it is AI rather than any specific measurable defect.", "body_md": "Have you ever eaten a dish where the flavor was all there but the texture was just… off? Maybe it’s perfectly seasoned, but the mouthfeel is gritty or somehow wrong. Texture matters for good food.\n\nThat’s often my feeling when reading completely AI-generated content, even if the underlying content is decent. I wrote a prior piece [criticizing a trend where writers would cite ChatGPT or Claude to implicitly offset blame if their research was wrong](https://weightythoughts.com/p/please-stop-citing-chatgpt-and-claude). This isn’t that.\n\nAn experienced and prolific author recently described to me the process of composing words as similar to composing music. There’s a rhythm and flow to it. I agree. This “rhythm” is hard to quantitatively measure. But we tend to know it when we see it.\n\nHumans are very good at spotting “unnatural” things, especially those that imitate humans. This is the entire origin of the uncanny valley, where hyperrealistic robots or animations mimicking humans feel off-putting to us, and we can intuitively catch the tiniest mismatches. In music, producers often need to microshift their timings to make it deliberately less perfectly on-beat—like there would be in human performance—to prevent it from feeling robotic.\n\n## What is it about AI writing exactly?\n\nIn line with societal angst about AI, especially in writing and publishing circles, people talk about how much they hate AI writing. Pangram’s rollout on Substack has created mini-witch hunts, and influencers like [Hank Green](https://www.vanityfair.com/story/hank-green-ai-youtube) who are accused of using AI face being “cancelled.”\n\nNone of this actually describes *what* is wrong with AI writing, though. Being AI is sufficient to hate it.\n\nAs longtime readers know, that’s not my stance. Although I didn’t use AI at all in my book (publisher requirement, plus it was too early in 2022 and 2023 to have AI be useful even as a research assistant), [I’ve been transparent about AI use before it was cool and tried to see how far I can push it in my Substack](https://weightythoughts.com/p/is-ai-in-writing-a-problem). \n\nThe answer to its usefulness is quite mixed. It never gets anywhere close to passing the bar on its own and basically gets totally rewritten—enough so that even when I use AI, usually Claude, to do a first draft, it never comes up as AI writing on Pangram. It’s not bad as a carefully monitored research assistant ([especially with a lot of adversarial reviews of its work!](https://weightythoughts.com/p/stop-wasting-human-time-with-ai-mistakeshallucin)).\n\n*As a side note: I pretty much entirely stopped having AI try its hand at a first draft. I always ended up changing it so much that, for now, it is far easier just to do it myself.*\n\n[Pangram](https://open.substack.com/users/524219152-pangram?utm_source=mentions), love it or hate it, is interesting in part because while we often know AI writing when we see it, it isn’t really that clear—quantitatively speaking—what makes writing AI.\n\n### What it is not\n\nThe typical hallmarks of AI, like the oft-complained-about em dash—which I will never give up—tend to be poor indicators. That should make sense. If something is so obvious, the next model version will inevitably remove it. For example, recent Claude models have made “load-bearing” a punchline given how much they seem to love it. That’s inevitably going to disappear, if not naturally, then definitely through post-training.\n\nAdditionally, you get a lot of variation just by using different models. They don’t all write the same. While they all do feel like AI to me, there’s no static characteristic that’s simple to measure. That is, of course, [Pangram](https://open.substack.com/users/524219152-pangram?utm_source=mentions)’s entire claim to fame. Whether or not you believe them, I do expect that they’ll need to continuously train their classifiers because models will keep changing their habits.\n\nMusic is a little ahead of us here, mostly because audio is easier to measure than prose (no offense to music, but notes don’t have explicit connotation and denotation when stuck together, even if it is a “universal” language in a way!).\n\n[Researchers tested a commercial AI music detector across 30,000 tracks](https://doi.org/10.5334/tismir.254). The detectors worked almost perfectly on the generators they were trained on. Then they tried the detectors on a generator they hadn’t seen before. The result: 3 of 50 AI tracks were caught. Oh, but it gets worse!\n\nFunny enough, a bit of mangling also threw off the detectors. They simply lowered the sample rate on one of the generators’ tracks (Suno’s), which the detector had been trained on… and the detector caught none of the tracks. The detector was basing its verdict on encoding artifacts rather than anything musical.\n\nAlthough I don’t know what’s under Pangram’s hood, I’d expect they need a pretty robust retraining loop just to keep up, given how much models change from release to release.\n\n### What it is—at least for now\n\nAgain, it’s really, really hard to *precisely* say what creates the AI writing uncanny valley. I expect that whatever metrics one comes up with will expire in a year or two—either simply by model progress or by deliberate action from the AI labs.\n\nNevertheless, we do have *some* ability to measure why AI writing seems to lack the unique rhythmic quality of human writing.\n\nWhile I did say that normal “AI tells” are unreliable, certain models make themselves quite obvious. If I said “load-bearing” everywhere in this piece, even if I wrote it myself, a daily user of Claude would likely have PTSD flashbacks. As it stands, even with a fairly long style guide, I found that Claude consistently had stylistic tics that I would immediately rip out of its drafts.\n\nUsually, the reason is this kind of short sentence—“mic drop” (as I call it)—is meant to emphasize the climax of an argument. Claude, in particular, loves to put mic drops *everywhere*. And, of course, if everything is emphasized, nothing is emphasized.\n\nThis has pretty obvious parallels in both storytelling and music. You need to have moments where you pull back, create tension, and vary things up. Constantly hitting the crescendo climax non-stop just means that you’re sonic noise and an ear-sore.\n\nNevertheless, a lot of us experience this particular tic through Claude mostly because Claude is what we use. It isn’t exclusive to it, though—models do vary here. Every model has its own habits.\n\nBut broadly speaking, all of them tend to have similar “rhythmic” issues stemming from the fact that they are token generators. For the same reason there’s [context rot—where LLMs get dumber the longer you chat with them](https://weightythoughts.com/p/ai-dementiawhy-your-agent-gets-progressively)—LLMs inevitably have limitations in being able to perfectly track long writing in the way a human would.\n\nWe’ve seen some studies that basically show this. LLMs tend to converge on similar arguments. After all, their training set is similar because they all just ate the internet. Additionally, their sentences, as one would expect, are templates which also all resemble one another more than human sentences do. LLMs, as it says in their very name, are language models and you… well, model… sentences in language.\n\nAgain, this is a “point in time” metric. I have zero doubt that if *this* ends up being the main hallmark of AI writing, the labs will find a way to “jitter” sentences to “break” these measurements. It’s like benchmaxxing—once someone creates a benchmark, inevitably the labs will have releases that “game” those metrics.\n\nAs an aside, this is why *certain* models (say… like Gemini) seem to top benchmarks… but are terrible to use in actual practice.\n\nOf course, breaking up this weird templating will likely help make the text feel more human, benchmaxxing or not. Going back to music, it’s possible to have musical notes hit *exactly* at the right timing. This is called quantization. Many music producers, even in electronic music, add in “swing” or simply hand-play certain parts to add the right “imperfection” and make their pieces feel more human. Similarly, less “quantized” AI text will probably read as more human. But, again, like music, there’s still a gap between just random jitter and *true* human performance.\n\nIt comes down to semantics. There’s a *reason* why certain rhythms are adopted for certain emphases and emotional impact. Just as true AI or algorithmically generated music often feels “weird” to listeners because the *meaning* or *intent* doesn’t seem to be there, I do expect that ineffable quality will be a gap for a long, long time in AI writing.\n\nIn fact, without more fundamental changes in how the models work, I don’t think that gap will be closed at all. While “embedding” (translating tokens into numerical vectors… which also creates numerical associations between concepts) kind of gets at semantics, it’s not really the same thing as composing an overall argument using language and the flow of a writing piece in mind.\n\n## Why Does it Matter?\n\nTo some degree, I’m in the camp that I don’t care if you use AI if I can’t notice it. I say “to some degree,” perhaps because, as a writer, I’m more sensitive to the “musical” rhythm of good writing.\n\nIf you use AI to do research, so long as you weren’t lazy and did the equivalent of “Let me Google That For You” and added something I can trust to my understanding, I’m happy enough. It doesn’t matter if you used an encyclopedia, Google, or ChatGPT. You put your stamp on it, which means you stake your reputation on it, which, in turn, allows me to trust what you put out. It literally is the original purpose and meaning behind “brands” in adding a mark of trustworthiness.\n\nIf you draft the piece with AI… but do such extensive edits—similar to my own process with Substack for a while—that it’s unrecognizable as AI, I also don’t really care. You basically inserted enough “human performance” that everything I care about, even “musically,” is present anyway.\n\nAdmittedly, given my experience, I’m not sure a first AI draft actually saves time with models as they are today… but that’s irrelevant to the end result. Whether you waste time doing it that way or not is beside the point.\n\nNo matter what, if I’m reading *you* as an author, I want it to be *you* as the author. You need to add what makes you you. Otherwise, I could just go to Claude or ChatGPT and just go prompt it myself.\n\n### Who is the Author, Exactly?\n\nThe concept of authorship is something I wrangled with in my book. My conclusion, between *[Death of an Author](https://en.wikipedia.org/wiki/Death_of_an_Author_%28novella%29)* (a totally “ChatGPT written” book by @Stephen Marche), pop artists like Andy Warhol often using assistants to make the “actual art,” and the very art form of photography itself… authorship or creative ownership is more about the idea than the tool used to execute it—including AI.\n\n*Death of an Author* is an example I especially love, since the author kept regenerating it until he was satisfied with what he got (and, besides that, wrote the plot himself). In an extreme case, if you “type” your book by hitting regenerate millions of times until it looks like what you would write—or at least what passes your quality bar—is that *really* that much different than just opening Microsoft Word and typing it? Wasteful/inefficient, sure, but it’s the same end result.\n\nThere are less extreme thought experiments, but I think most of them intuitively come out on the side of “it’s equivalent to typing” if you think hard enough about it. While there is a spectrum, if the human is the arbiter of quality and taste and lets it go out into the world having satisfied that bar… well, again, that’s why I follow specific authors anyway.\n\n### AI Watermarks\n\nThis particular point has some recent news salience thanks to Anthropic. Recently, in response to EU regulations (which, as is typical, created many more unintended bad consequences than things it solved…), Anthropic declared that they’d be using watermarks for AI writing going forward. The relevant AI Act provisions took effect on August 2nd, and per [Anthropic’s own documentation](https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content), models launched on or after that date carry marks at launch. While they only needed to do it within the EU, they rolled it out worldwide.\n\nNow, what’s bad about this? Well, two things.\n\nOne is that watermarks, in theory, will degrade the quality of writing. The way you can have “invisible” watermarks is by deliberately introducing certain non-randomness in word choice (well, token generation, but you get what I mean) that can be picked up. [Google’s own research on the technique](https://www.nature.com/articles/s41586-024-08025-4) concedes it causes “some reduction to inter-response diversity.”\n\nI just spent the entire piece talking about this kind of textual or rhythm quality causing AI generated writing to feel “off.” The proof will be in the pudding, but it’s hard to imagine this watermark not moving Claude in the wrong direction in “good writing.” But my reaction to that is somewhat of a shoulder shrug. I expect if it makes Claude’s writing feel *really bad*, Anthropic will change it. It’s not like I don’t *already* have to massively edit for word choice and rhythm anyway, so making it slightly worse is somewhat irrelevant to my use case.\n\nThe second is worse and has to do with “authorship.” Using their models, even if you’re doing a copywriting pass purely, you can have Claude “claim ownership” using watermarks. I think it’s silly in general, but certain other authors like [Ben Thompson](https://stratechery.com/2026/anthropics-watermarking-how-it-probably-works-worse-than-it-seems/) and [John Gruber](https://daringfireball.net/linked/2026/08/27/the-load-bearing-vocabulary-of-claude) have been fairly pissed off by the entire thing.\n\nAfter all, as per [Anthropic](https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content) itself:\n\n*Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source*\n\nIn [Dithering](https://dithering.passport.online/member/episode/watermarks), Ben Thompson asserts that this is conveniently aligned with Anthropic’s view that Claude is basically its own entity and deserves credit for everything it outputs. After all, once the entire world is using Claude, all of it should partly belong to it—and Anthropic. \n\nThis is somewhat typical [AI lab arrogance](https://weightythoughts.com/p/when-genius-failsthe-intellectual). It is certainly convenient for Anthropic to use the EU’s regulation to try to put its mark on everything.\n\n### In the Longer Term\n\nPerhaps the watermark only matters insofar as the public still has extreme reactions to AI writing. After all, if every single word processor in the world used an “AI copyeditor” like Claude that put in small watermarks, its presence would be entirely irrelevant.\n\n**I do think we’ll eventually go toward a world in the future where AI use is as incidental as computer use.** Basically, you’d find a new employee who comes up to you and says, “I don’t use AI,” just as weird as one today who comes up to you and says, “Well, I don’t really do ‘computers’” and hands you a stack of handwritten notes… which you’d likely then need to transcribe into a computer.\n\nIt’s part of the reason I’m not really that fussed about the watermark, even if I do think it’s misguided.\n\nBut putting aside practical considerations, there’s a deeper philosophical reason I end up where I am. And that is that AI doesn’t replace humans, at least not with the AI that we have today. It has different strengths and weaknesses… and especially for matters of taste, I still care about the human behind the artistic piece, whether it’s visual, audio, or written. That *really* shouldn’t be surprising. After all, there’s no necessarily “objective” reason why a human-performed music piece is better than a jittered or quantized piece.\n\nIf AI can do a lot of the practical stuff better and improve productivity in the world, I’m all for it. That’s been the way progress works, from shovels to steam engines to nuclear-powered ships. None of these obsoleted humans; they simply changed the nature of work.\n\nIf you think, from the creative side, that AI will simply wipe out humans… well, perhaps you have less faith in humanity than I do.\n\nBut even if you do have less faith, I think you should logically expect that humans—by matters of *pure taste*—will prefer things with a human touch… because we are human and AI is not.\n\n# **Thanks for reading!**\n\nI hope you enjoyed this post. If you’d like to learn more about AI’s past, present, and future in an easy-to-understand way, I’ve published a book titled *What You Need to Know About AI*.\n\nYou can order the book on [Amazon](https://amzn.to/4qCERMX), [Barnes & Noble](https://bit.ly/barnesandnobleaibook), [Bookshop](https://bit.ly/bookshopaibook), or pick up a copy in-person at a [local bookstore](https://www.smartaibook.com/buy).", "url": "https://wpnews.pro/news/why-is-ai-bad-at-writing", "canonical_source": "https://weightythoughts.com/p/why-is-ai-bad-at-writing", "published_at": "2026-09-08 13:03:18+00:00", "updated_at": "2026-09-22 13:53:57.015929+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "generative-ai", "ai-tools"], "entities": ["Claude", "Pangram", "Substack", "Hank Green", "Weighty Thoughts"], "alternates": {"html": "https://wpnews.pro/news/why-is-ai-bad-at-writing", "markdown": "https://wpnews.pro/news/why-is-ai-bad-at-writing.md", "text": "https://wpnews.pro/news/why-is-ai-bad-at-writing.txt", "jsonld": "https://wpnews.pro/news/why-is-ai-bad-at-writing.jsonld"}}