{"slug": "something-bugs-me-about-agi-ai-llm-what-if-we-back-paddled-1000-year", "title": "Something bugs me about AGI AI LLM, what if we back paddled 1000 year", "summary": "A developer raises a hypothesis that increasingly capable LLM coding agents could make small software failures so cheap to repair that they stop serving as visible signals of deeper structural problems, potentially shifting risk from frequent minor incidents toward rarer but more correlated catastrophic failures. The engineer argues the concern is not that AI writes bad code, but that it may become extremely good at continuously patching symptoms of structural weakness, and calls for the hypothesis to be tested rather than assumed.", "body_md": "# What If AI Makes the Cracks Too Cheap to Notice?\n\n### A monkey-business question about LLMs, tail risk, human disagreement, yin and yang, Chinese medicine, and why making every small failure cheap might make the largest failures more expensive\n\nI have a question about AI coding that I cannot get out of my head.\n\nI don’t know whether the hypothesis is right.\n\nI am not claiming I have discovered some hidden law of AI, economics, software engineering, or human civilization. I am trying to point at something that smells strange to me and ask whether anyone has seriously measured it.\n\nThe current story goes something like this:\n\n**LLMs can write code.**\n\nLLMs can review code.\n\nLLMs can debug code.\n\nAgents can fix tests, refactor repositories, investigate failures, write documentation, and increasingly perform longer sequences of software work.\n\nEventually we start talking as if we have moved one abstraction layer upward again:\n\n**assembly → higher-level languages → natural language.**\n\nMaybe.\n\nBut something about that framing bothers me.\n\nNot because I think LLM coding is fake.\n\nAlmost the opposite.\n\n**What if it works well enough locally that it changes which structural failures remain visible?**\n\n# The Building That Repairs Its Own Cracks\n\nImagine a large building.\n\nHistorically, little cracks appearing in the walls were annoying.\n\nSomeone had to notice the crack.\n\nSomeone had to investigate it.\n\nSomeone had to figure out whether it was cosmetic or structural.\n\nSomeone had to repair it.\n\nAll of this consumed scarce human attention.\n\nThat sucked.\n\nNow imagine that we invent an absurdly capable Crack Repair Machine™.\n\nA crack appears.\n\n**ZZZZT.**\n\nFixed.\n\nAnother crack appears.\n\n**ZZZZT.**\n\nFixed.\n\nTwenty cracks?\n\nWho cares?\n\nThe machine can fix twenty cracks before lunch.\n\nThis sounds strictly better.\n\nAnd perhaps it is.\n\nBut consider another possibility.\n\nWhat if some of those annoying little cracks were also **weak signals of load moving through the building incorrectly?**\n\nPreviously, because cracks were expensive, humans occasionally had to ask:\n\nWhy the hell does this wall keep cracking?\n\nNow the local repair cost approaches zero.\n\nSo instead we get:\n\n`CRACK → PATCH → NEXT`\n\n`CRACK → PATCH → NEXT`\n\n`CRACK → PATCH → NEXT`\n\nThe building looks fantastic.\n\nUntil one day the problem isn’t a crack.\n\nThe load-bearing structure has moved.\n\n# Cheap Local Failure Does Not Necessarily Mean Lower Global Risk\n\nThis is the risk-distribution question I actually care about.\n\nSuppose the old world produces lots of small failures.\n\nThey are visible.\n\nThey are expensive.\n\nHumans hate them.\n\nBut catastrophic failure is relatively rare because the small failures continuously expose weaknesses in the structure.\n\nNow introduce machinery that makes local failure extraordinarily cheap:\n\n- generate code cheaply\n- regenerate broken code cheaply\n- repair tests cheaply\n- patch failures cheaply\n- rewrite implementations cheaply\n- replace components cheaply\n- produce enormous amounts of new code cheaply\n\nWonderful.\n\nBut what happened to the distribution of risk?\n\nPerhaps we reduced the probability and cost of small incidents.\n\nBut did we also change the probability of the tail?\n\nCould we accidentally move from:\n\n**many visible small failures + rare catastrophe**\n\ntoward:\n\n**almost invisible small failures + rarer-looking but more correlated catastrophe?**\n\nThe scary version isn’t:\n\n“AI writes bad code.”\n\nThat’s boring.\n\nThe scary version is:\n\n**AI becomes extremely good at continuously repairing symptoms of structural weakness, thereby reducing the human attention those symptoms previously attracted.**\n\nMaybe that hypothesis is completely wrong.\n\nGreat.\n\n**Test it.**\n\n# Now Make the Problem Worse: What If Everyone Uses the Same Yapper?\n\nThis is where my monkey brain wanders away from software engineering.\n\nImagine 100 humans.\n\nThose 100 humans do not have one objective function.\n\nThey have something closer to:\n\nG1,G2,G3,…,G100G_1, G_2, G_3,\\ldots,G_{100}\n\nOne wants money.\n\nOne wants stability.\n\nOne wants to make beautiful things.\n\nOne wants to go home.\n\nOne wants status.\n\nOne wants his children safe.\n\nOne wants to play Brood War.\n\nOne wants to protect the customer.\n\nOne wants to get promoted.\n\nOne thinks the entire project is stupid.\n\nOne wants bananas.\n\nExcellent.\n\nHumanity is an absolutely disgusting distributed system.\n\nAnd perhaps that is a feature.\n\nEvery human has some amount of time, attention, energy, money, authority, reputation, labor and physical action.\n\nEvery day they allocate those resources according to their own weird objective function.\n\nSo everybody effectively gets an invisible vote.\n\nNot a political vote.\n\nA **resource-allocation vote**:\n\nI will spend one hour on this and zero hours on that.\n\nMultiply that across billions of people and perhaps what we call a market, organization, culture, society—or civilization—is partly the emergent equilibrium produced by billions of incompatible objective functions continuously pushing against each other.\n\nThat friction looks inefficient.\n\nBut what if the friction is carrying structural information?\n\n# What Happens When the Yapper Does Most of the Work?\n\nNow give all 100 humans an LLM.\n\nEventually, perhaps, give the LLMs increasingly large portions of the work.\n\nWhat happens to those 100 unique vectors?\n\nDo we preserve:\n\n{G1,G2,…,G100}\\{G_1,G_2,\\ldots,G_{100}\\}\n\nor do layers of shared models, shared post-training, shared interfaces, shared optimization patterns, shared agent frameworks and shared abstractions begin projecting those vectors into some smaller effective space?\n\nI don’t know.\n\nBut I desperately want somebody to measure it.\n\nBecause if the machinery performing more of humanity’s work reduces the behavioral diversity through which those 100 different goals previously expressed themselves, then this isn’t merely a productivity question.\n\nIt becomes a **stability question**.\n\nThe system may become locally more efficient while losing some of the messy counterweights that kept it globally balanced.\n\nThat thought sent me somewhere ridiculous.\n\nIt sent me back to ancient Chinese clichés.\n\n# Yin, Yang, Harmony—and the Most Criminal Documentation Strategy Ever Invented\n\n阴阳。\n\n和谐。\n\n八卦。\n\nQi.\n\nBalance.\n\nFlow.\n\nThese words are so overused that they can become almost meaningless.\n\nBut lately I have been wondering whether there is an interesting information problem hidden underneath them.\n\nImagine generations of humans observing incredibly complicated systems:\n\nweather,\n\nfood,\n\nillness,\n\nfamilies,\n\npolitics,\n\nwar,\n\nagriculture,\n\nthe body,\n\nemotion,\n\nsocial relationships,\n\npower.\n\nThey do not have modern instrumentation.\n\nThey do not have our mathematical vocabulary.\n\nThey do not have databases containing every intermediate observation.\n\nThey live inside the system and repeatedly experience it.\n\nEventually somebody develops an extremely compressed intuition:\n\nThere is an imbalance here.\n\nAnd then commits the greatest documentation crime imaginable.\n\nInstead of leaving us the complete reasoning trace, telemetry, training corpus, failed hypotheses and state machine, he writes:\n\n**阴阳。**\n\nBRO.\n\nWHERE ARE THE FUCKING LOGS?\n\n😂\n\nIt is almost like some ancient expert spent 70 years training an internal anomaly detector and then shipped the model weights without the training data.\n\n# The Walking Model\n\nThis is what fascinates me about the archetype of the old master.\n\nThe master doesn’t necessarily possess a convenient explicit decision tree:\n\n`IF A AND B THEN C`\n\nInstead, the person may have spent decades exposing a biological neural system to outcomes.\n\nSomething looks wrong.\n\nSomething feels wrong.\n\nSomething sounds wrong.\n\nThey inspect.\n\nSometimes they are wrong.\n\nSometimes they are right.\n\nFeedback arrives.\n\nWeights update.\n\nRepeat for 50 years.\n\nEventually:\n\n`WORLD STATE → WTF CHECK THAT`\n\nfires before the person can fully verbalize why.\n\nWe call that:\n\n**experience.**\n\n**intuition.**\n\n**sixth sense.**\n\n**expert judgment.**\n\nMaybe sometimes it is wisdom.\n\nMaybe sometimes it is complete bullshit.\n\nThat’s exactly why I want the logs.\n\n# 老中医.exe\n\nNow take the stereotypical old Chinese medicine practitioner.\n\nIgnore for a moment the question of which specific treatments scientifically work. That’s a separate empirical question.\n\nI’m interested in the **information architecture of the practitioner**.\n\nThe practitioner observes:\n\npulse,\n\nskin,\n\nvoice,\n\ntemperature,\n\nsleep,\n\nappetite,\n\npain,\n\nenergy,\n\nhistory,\n\nmovement,\n\nwhatever else their tradition tells them to inspect.\n\nAfter enough experience, the practitioner may effectively become a walking pattern-recognition system.\n\nThen the student asks:\n\nWhy?\n\nAnd history answers:\n\nSomething something qi.\n\nFUCK.\n\n😂\n\nThe experienced practitioner may contain an enormous compressed model produced by decades of observations, while the transferable documentation contains only fragments of the reasoning process.\n\nSo the next generation has to partly reconstruct the model by **living through another enormous training run**.\n\nThat is simultaneously beautiful and horrifying engineering.\n\n# Western Mass Production vs. The Old Guy Who “Feels the Qi”\n\nThis gave me another stupid analogy.\n\nImagine two approaches to a difficult target.\n\n### Approach A: Mass-produced artillery\n\nWe have a treatment/intervention that:\n\n- is standardized\n- is documented\n- is trainable at scale\n- works often enough\n- has known procedures\n- can be deployed by many practitioners\n\nIt may not perfectly model the individual target.\n\nBut we can manufacture lots of ammunition.\n\nSo:\n\n**THROW ROCK → observe → adjust → throw next rock.**\n\nThis is enormously valuable.\n\nCivilization needs scalable solutions.\n\n### Approach B: Precision-guided old-turtle missile\n\nThen somewhere there is a ridiculous domain expert who looks at the same target and says:\n\nDon’t shoot there.\n\nShoot **there**.\n\nOne tiny intervention.\n\nHuge effect.\n\nEveryone asks:\n\nHOW THE FUCK DID YOU KNOW?\n\nAnd the answer is:\n\nForty years.\n\nFantastic.\n\nCompletely unscalable API.\n\n# Depression, “Go Take a Walk,” and Cheap Rocks\n\nThis analogy becomes especially funny around human problems.\n\nSomeone is stuck.\n\nNot metaphorically “lazy.”\n\nTheir system is not producing enough forward movement.\n\nMaybe motivation is low.\n\nMaybe reward feels absent.\n\nMaybe everything looks expensive.\n\nMaybe some unresolved thought continuously reacquires attention.\n\nMaybe sleep is destroyed.\n\nMaybe the problem is biological, social, financial, emotional, environmental—or seventeen things simultaneously.\n\nThe full state space is enormous.\n\nAnd somebody says:\n\n**Go take a walk.**\n\nThis can sound insultingly stupid.\n\nBut from another perspective, it is a fascinating mass-produced rock.\n\nWalking is:\n\n- cheap\n- accessible to many people\n- relatively low-risk for many people\n- physically activating\n- environmentally changing\n- attention-shifting\n- easy to prescribe\n- easy to test\n\nIt does not mean:\n\nWALKING SOLVES DEPRESSION.\n\nIt means:\n\nHere is one inexpensive intervention that sometimes perturbs a stuck system in a useful direction.\n\nThrow the rock.\n\nObserve.\n\nIf nothing useful happens, don’t worship the rock.\n\nTry to understand the system better.\n\n# And Suddenly I’m Back to LLMs\n\nThis is where my wall.txt somehow loops back to the beginning.\n\nWhat are LLMs doing to our civilization’s intervention economics?\n\nThey make certain classes of rocks **absurdly cheap**.\n\nNeed code?\n\nThrow rock.\n\nNeed rewrite?\n\nThrow rock.\n\nNeed analysis?\n\nThrow rock.\n\nNeed summary?\n\nThrow rock.\n\nNeed debugging?\n\nThrow rock.\n\nNeed ten hypotheses?\n\nThrow ten rocks.\n\nNeed one thousand?\n\nFuck it.\n\nThe marginal cost keeps collapsing.\n\nThat is extraordinary.\n\nBut if throwing rocks becomes essentially free, **sensing becomes more important, not less.**\n\nBecause now the limiting resource isn’t necessarily:\n\nCan we produce an intervention?\n\nThe limiting resource becomes:\n\n**Do we know where to aim?**\n\nAnd perhaps even:\n\n**Can we still detect that the underlying structure has moved when the machine continuously repairs every cheap local symptom?**\n\n# The AI Era Might Be a Sensing Problem Disguised as a Generation Problem\n\nThis is the hypothesis I currently find interesting.\n\nWe are obsessed with generation because generation is visibly improving.\n\nMore code.\n\nMore text.\n\nMore agents.\n\nMore actions.\n\nMore automation.\n\nMore rocks.\n\nBut perhaps the scarce resource moves somewhere else.\n\nToward:\n\n**sensing.**\n\n**indexing.**\n\n**anomaly detection.**\n\n**goal preservation.**\n\n**knowing which bit matters.**\n\n**knowing whose objective function is being optimized.**\n\n**knowing when not to act.**\n\n**knowing when a tiny crack is actually telling you that the building moved.**\n\nIf output becomes cheap enough, then the person/system capable of saying:\n\n**STOP. LOOK AT THIS ONE BIT.**\n\nmay become disproportionately valuable.\n\n# The 100-Monkey Question\n\nSo here is my actual war cry.\n\nImagine 100 humans originally performing 100 pieces of work.\n\nEach carries a different history, goal, incentive, fear, preference, intuition and definition of value.\n\nNow LLM systems perform 80% of their intermediate cognitive work.\n\nWhat happens to the diversity of the resulting decisions?\n\nDo the 100 humans remain 100 independent vectors?\n\nOr does shared machinery begin introducing correlated blind spots?\n\nDoes local productivity increase while systemic diversity decreases?\n\nDo tiny errors become cheaper while catastrophic correlated errors become more probable?\n\nDoes disagreement disappear from intermediate reasoning because the same machinery increasingly mediates everybody’s thought-to-action pipeline?\n\nDoes the system become more harmonious?\n\nOr does it only **look** harmonious because we accidentally removed some of the sensors that previously expressed imbalance?\n\nI don’t know.\n\nThat’s why I want the wind tunnel.\n\n# Give Me the Logs\n\nI don’t want:\n\nAI good.\n\nI don’t want:\n\nAI bad.\n\nI don’t particularly care whether somebody calls it AGI.\n\nGive me the system.\n\nGive me the objective.\n\nGive me the environment.\n\nGive me the constraints.\n\nGive me the human.\n\nGive me the machine.\n\nGive me the information available to each.\n\nGive me the timestamps.\n\nGive me the failures.\n\nGive me the little cracks.\n\nThen let’s change one thing and see what moves.\n\nMaybe my entire hypothesis gets flushed down the toilet.\n\nExcellent.\n\nBut I increasingly suspect that the interesting question of the LLM era isn’t merely:\n\n**How much work can the machine do?**\n\nIt may be:\n\n**When the machine makes action and repair almost free, what becomes expensive—and which weak signals stop receiving human attention because we no longer need humans to deal with the small failures?**\n\nAnd one level above that:\n\n**If billions of independently weird humans were part of the stabilizing feedback mechanism of civilization, what happens when increasingly large portions of their actions are mediated by a smaller family of machines?**\n\nMaybe nothing.\n\nMaybe productivity goes up and everything is wonderful.\n\nMaybe new diversity appears.\n\nMaybe humans simply move their attention one abstraction layer upward.\n\nOr maybe we discover that all the stupid friction, disagreement, duplicated work, weird intuition and incompatible goals were carrying information we didn’t know we were using.\n\nI genuinely don’t know.\n\nBut if we’re about to throw infinite cheap rocks—\n\n**I would really like to know whether we’re getting better at sensing where the fuck to aim.**\n\n*我了个米饭。*\n\n**GG.**", "url": "https://wpnews.pro/news/something-bugs-me-about-agi-ai-llm-what-if-we-back-paddled-1000-year", "canonical_source": "https://shatteringtheabyss.substack.com/p/llm-agi-god-congratulations-humanity", "published_at": "2026-09-22 02:10:04+00:00", "updated_at": "2026-09-22 02:23:23.077953+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-safety", "ai-research", "developer-tools"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/something-bugs-me-about-agi-ai-llm-what-if-we-back-paddled-1000-year", "markdown": "https://wpnews.pro/news/something-bugs-me-about-agi-ai-llm-what-if-we-back-paddled-1000-year.md", "text": "https://wpnews.pro/news/something-bugs-me-about-agi-ai-llm-what-if-we-back-paddled-1000-year.txt", "jsonld": "https://wpnews.pro/news/something-bugs-me-about-agi-ai-llm-what-if-we-back-paddled-1000-year.jsonld"}}