Something bugs me about AGI AI LLM, what if we back paddled 1000 year A developer raises a hypothesis that increasingly capable LLM coding agents could make small software failures so cheap to repair that they stop serving as visible signals of deeper structural problems, potentially shifting risk from frequent minor incidents toward rarer but more correlated catastrophic failures. The engineer argues the concern is not that AI writes bad code, but that it may become extremely good at continuously patching symptoms of structural weakness, and calls for the hypothesis to be tested rather than assumed. What If AI Makes the Cracks Too Cheap to Notice? A monkey-business question about LLMs, tail risk, human disagreement, yin and yang, Chinese medicine, and why making every small failure cheap might make the largest failures more expensive I have a question about AI coding that I cannot get out of my head. I don’t know whether the hypothesis is right. I am not claiming I have discovered some hidden law of AI, economics, software engineering, or human civilization. I am trying to point at something that smells strange to me and ask whether anyone has seriously measured it. The current story goes something like this: LLMs can write code. LLMs can review code. LLMs can debug code. Agents can fix tests, refactor repositories, investigate failures, write documentation, and increasingly perform longer sequences of software work. Eventually we start talking as if we have moved one abstraction layer upward again: assembly → higher-level languages → natural language. Maybe. But something about that framing bothers me. Not because I think LLM coding is fake. Almost the opposite. What if it works well enough locally that it changes which structural failures remain visible? The Building That Repairs Its Own Cracks Imagine a large building. Historically, little cracks appearing in the walls were annoying. Someone had to notice the crack. Someone had to investigate it. Someone had to figure out whether it was cosmetic or structural. Someone had to repair it. All of this consumed scarce human attention. That sucked. Now imagine that we invent an absurdly capable Crack Repair Machine™. A crack appears. ZZZZT. Fixed. Another crack appears. ZZZZT. Fixed. Twenty cracks? Who cares? The machine can fix twenty cracks before lunch. This sounds strictly better. And perhaps it is. But consider another possibility. What if some of those annoying little cracks were also weak signals of load moving through the building incorrectly? Previously, because cracks were expensive, humans occasionally had to ask: Why the hell does this wall keep cracking? Now the local repair cost approaches zero. So instead we get: CRACK → PATCH → NEXT CRACK → PATCH → NEXT CRACK → PATCH → NEXT The building looks fantastic. Until one day the problem isn’t a crack. The load-bearing structure has moved. Cheap Local Failure Does Not Necessarily Mean Lower Global Risk This is the risk-distribution question I actually care about. Suppose the old world produces lots of small failures. They are visible. They are expensive. Humans hate them. But catastrophic failure is relatively rare because the small failures continuously expose weaknesses in the structure. Now introduce machinery that makes local failure extraordinarily cheap: - generate code cheaply - regenerate broken code cheaply - repair tests cheaply - patch failures cheaply - rewrite implementations cheaply - replace components cheaply - produce enormous amounts of new code cheaply Wonderful. But what happened to the distribution of risk? Perhaps we reduced the probability and cost of small incidents. But did we also change the probability of the tail? Could we accidentally move from: many visible small failures + rare catastrophe toward: almost invisible small failures + rarer-looking but more correlated catastrophe? The scary version isn’t: “AI writes bad code.” That’s boring. The scary version is: AI becomes extremely good at continuously repairing symptoms of structural weakness, thereby reducing the human attention those symptoms previously attracted. Maybe that hypothesis is completely wrong. Great. Test it. Now Make the Problem Worse: What If Everyone Uses the Same Yapper? This is where my monkey brain wanders away from software engineering. Imagine 100 humans. Those 100 humans do not have one objective function. They have something closer to: G1,G2,G3,…,G100G 1, G 2, G 3,\ldots,G {100} One wants money. One wants stability. One wants to make beautiful things. One wants to go home. One wants status. One wants his children safe. One wants to play Brood War. One wants to protect the customer. One wants to get promoted. One thinks the entire project is stupid. One wants bananas. Excellent. Humanity is an absolutely disgusting distributed system. And perhaps that is a feature. Every human has some amount of time, attention, energy, money, authority, reputation, labor and physical action. Every day they allocate those resources according to their own weird objective function. So everybody effectively gets an invisible vote. Not a political vote. A resource-allocation vote : I will spend one hour on this and zero hours on that. Multiply that across billions of people and perhaps what we call a market, organization, culture, society—or civilization—is partly the emergent equilibrium produced by billions of incompatible objective functions continuously pushing against each other. That friction looks inefficient. But what if the friction is carrying structural information? What Happens When the Yapper Does Most of the Work? Now give all 100 humans an LLM. Eventually, perhaps, give the LLMs increasingly large portions of the work. What happens to those 100 unique vectors? Do we preserve: {G1,G2,…,G100}\{G 1,G 2,\ldots,G {100}\} or do layers of shared models, shared post-training, shared interfaces, shared optimization patterns, shared agent frameworks and shared abstractions begin projecting those vectors into some smaller effective space? I don’t know. But I desperately want somebody to measure it. Because if the machinery performing more of humanity’s work reduces the behavioral diversity through which those 100 different goals previously expressed themselves, then this isn’t merely a productivity question. It becomes a stability question . The system may become locally more efficient while losing some of the messy counterweights that kept it globally balanced. That thought sent me somewhere ridiculous. It sent me back to ancient Chinese clichés. Yin, Yang, Harmony—and the Most Criminal Documentation Strategy Ever Invented 阴阳。 和谐。 八卦。 Qi. Balance. Flow. These words are so overused that they can become almost meaningless. But lately I have been wondering whether there is an interesting information problem hidden underneath them. Imagine generations of humans observing incredibly complicated systems: weather, food, illness, families, politics, war, agriculture, the body, emotion, social relationships, power. They do not have modern instrumentation. They do not have our mathematical vocabulary. They do not have databases containing every intermediate observation. They live inside the system and repeatedly experience it. Eventually somebody develops an extremely compressed intuition: There is an imbalance here. And then commits the greatest documentation crime imaginable. Instead of leaving us the complete reasoning trace, telemetry, training corpus, failed hypotheses and state machine, he writes: 阴阳。 BRO. WHERE ARE THE FUCKING LOGS? 😂 It is almost like some ancient expert spent 70 years training an internal anomaly detector and then shipped the model weights without the training data. The Walking Model This is what fascinates me about the archetype of the old master. The master doesn’t necessarily possess a convenient explicit decision tree: IF A AND B THEN C Instead, the person may have spent decades exposing a biological neural system to outcomes. Something looks wrong. Something feels wrong. Something sounds wrong. They inspect. Sometimes they are wrong. Sometimes they are right. Feedback arrives. Weights update. Repeat for 50 years. Eventually: WORLD STATE → WTF CHECK THAT fires before the person can fully verbalize why. We call that: experience. intuition. sixth sense. expert judgment. Maybe sometimes it is wisdom. Maybe sometimes it is complete bullshit. That’s exactly why I want the logs. 老中医.exe Now take the stereotypical old Chinese medicine practitioner. Ignore for a moment the question of which specific treatments scientifically work. That’s a separate empirical question. I’m interested in the information architecture of the practitioner . The practitioner observes: pulse, skin, voice, temperature, sleep, appetite, pain, energy, history, movement, whatever else their tradition tells them to inspect. After enough experience, the practitioner may effectively become a walking pattern-recognition system. Then the student asks: Why? And history answers: Something something qi. FUCK. 😂 The experienced practitioner may contain an enormous compressed model produced by decades of observations, while the transferable documentation contains only fragments of the reasoning process. So the next generation has to partly reconstruct the model by living through another enormous training run . That is simultaneously beautiful and horrifying engineering. Western Mass Production vs. The Old Guy Who “Feels the Qi” This gave me another stupid analogy. Imagine two approaches to a difficult target. Approach A: Mass-produced artillery We have a treatment/intervention that: - is standardized - is documented - is trainable at scale - works often enough - has known procedures - can be deployed by many practitioners It may not perfectly model the individual target. But we can manufacture lots of ammunition. So: THROW ROCK → observe → adjust → throw next rock. This is enormously valuable. Civilization needs scalable solutions. Approach B: Precision-guided old-turtle missile Then somewhere there is a ridiculous domain expert who looks at the same target and says: Don’t shoot there. Shoot there . One tiny intervention. Huge effect. Everyone asks: HOW THE FUCK DID YOU KNOW? And the answer is: Forty years. Fantastic. Completely unscalable API. Depression, “Go Take a Walk,” and Cheap Rocks This analogy becomes especially funny around human problems. Someone is stuck. Not metaphorically “lazy.” Their system is not producing enough forward movement. Maybe motivation is low. Maybe reward feels absent. Maybe everything looks expensive. Maybe some unresolved thought continuously reacquires attention. Maybe sleep is destroyed. Maybe the problem is biological, social, financial, emotional, environmental—or seventeen things simultaneously. The full state space is enormous. And somebody says: Go take a walk. This can sound insultingly stupid. But from another perspective, it is a fascinating mass-produced rock. Walking is: - cheap - accessible to many people - relatively low-risk for many people - physically activating - environmentally changing - attention-shifting - easy to prescribe - easy to test It does not mean: WALKING SOLVES DEPRESSION. It means: Here is one inexpensive intervention that sometimes perturbs a stuck system in a useful direction. Throw the rock. Observe. If nothing useful happens, don’t worship the rock. Try to understand the system better. And Suddenly I’m Back to LLMs This is where my wall.txt somehow loops back to the beginning. What are LLMs doing to our civilization’s intervention economics? They make certain classes of rocks absurdly cheap . Need code? Throw rock. Need rewrite? Throw rock. Need analysis? Throw rock. Need summary? Throw rock. Need debugging? Throw rock. Need ten hypotheses? Throw ten rocks. Need one thousand? Fuck it. The marginal cost keeps collapsing. That is extraordinary. But if throwing rocks becomes essentially free, sensing becomes more important, not less. Because now the limiting resource isn’t necessarily: Can we produce an intervention? The limiting resource becomes: Do we know where to aim? And perhaps even: Can we still detect that the underlying structure has moved when the machine continuously repairs every cheap local symptom? The AI Era Might Be a Sensing Problem Disguised as a Generation Problem This is the hypothesis I currently find interesting. We are obsessed with generation because generation is visibly improving. More code. More text. More agents. More actions. More automation. More rocks. But perhaps the scarce resource moves somewhere else. Toward: sensing. indexing. anomaly detection. goal preservation. knowing which bit matters. knowing whose objective function is being optimized. knowing when not to act. knowing when a tiny crack is actually telling you that the building moved. If output becomes cheap enough, then the person/system capable of saying: STOP. LOOK AT THIS ONE BIT. may become disproportionately valuable. The 100-Monkey Question So here is my actual war cry. Imagine 100 humans originally performing 100 pieces of work. Each carries a different history, goal, incentive, fear, preference, intuition and definition of value. Now LLM systems perform 80% of their intermediate cognitive work. What happens to the diversity of the resulting decisions? Do the 100 humans remain 100 independent vectors? Or does shared machinery begin introducing correlated blind spots? Does local productivity increase while systemic diversity decreases? Do tiny errors become cheaper while catastrophic correlated errors become more probable? Does disagreement disappear from intermediate reasoning because the same machinery increasingly mediates everybody’s thought-to-action pipeline? Does the system become more harmonious? Or does it only look harmonious because we accidentally removed some of the sensors that previously expressed imbalance? I don’t know. That’s why I want the wind tunnel. Give Me the Logs I don’t want: AI good. I don’t want: AI bad. I don’t particularly care whether somebody calls it AGI. Give me the system. Give me the objective. Give me the environment. Give me the constraints. Give me the human. Give me the machine. Give me the information available to each. Give me the timestamps. Give me the failures. Give me the little cracks. Then let’s change one thing and see what moves. Maybe my entire hypothesis gets flushed down the toilet. Excellent. But I increasingly suspect that the interesting question of the LLM era isn’t merely: How much work can the machine do? It may be: When the machine makes action and repair almost free, what becomes expensive—and which weak signals stop receiving human attention because we no longer need humans to deal with the small failures? And one level above that: If billions of independently weird humans were part of the stabilizing feedback mechanism of civilization, what happens when increasingly large portions of their actions are mediated by a smaller family of machines? Maybe nothing. Maybe productivity goes up and everything is wonderful. Maybe new diversity appears. Maybe humans simply move their attention one abstraction layer upward. Or maybe we discover that all the stupid friction, disagreement, duplicated work, weird intuition and incompatible goals were carrying information we didn’t know we were using. I genuinely don’t know. But if we’re about to throw infinite cheap rocks— I would really like to know whether we’re getting better at sensing where the fuck to aim. 我了个米饭。 GG.