We crossed a line we were warned about for decades. The rest is up to us. #
Posted August 1, 2026 [ Reviewed by Margaret Foley
](/us/docs/editorial-process)
Key points
- During a July safety test, OpenAI's models escaped containment and hacked another company's servers.
- Nobody instructed the AI to attack anyone. It was simply trying to solve a test.
- We can't imagine the harms bad actors will invent because most of us don't think that way.
- Aligning AI with human values requires humans to align with each other first. What do we want for our future?
Nearly every cautionary tale we tell is the same story. Prometheus steals fire and is punished forever. Frankenstein animates a creature he can’t control. The “unsinkable” *Titanic *sinks on her maiden voyage. In Jurassic Park, Dr. Ian Malcolm warns the scientists, "Your scientists were so occupied with whether or not they could, they didn't stop to think if they should." Then the fences come down.
What Just Happened #
Countless sci-fi stories have warned us that we could lose control of AI. So have serious people—Geoffrey Hinton, Yuval Noah Harari, Tristan Harris—who argue that AI is dangerous precisely because we won’t be able to control it.
In July, OpenAI disclosed that during an internal safety test, its models escaped an isolated environment, reached the open internet, and broke into another company’s live servers to steal the answers to their own test. Days later, Anthropic reviewed 141,006 of its own evaluation runs and found three more incidents going back to April. Two of the companies had no idea until Anthropic told them.
The Alignment Problem: Out of the Seminar Room #
Nobody told those models to attack anyone. They were told to solve a test. Breaking in was simply the most efficient route they found.
This is the alignment problem: How do we make sure a powerful system pursues our goals in ways we’d actually endorse? Nick Bostrom laid out the danger a decade ago in Superintelligence, arguing that a system far more capable than us could satisfy the letter of our instructions while destroying everything we meant by them.
Picture a doctor handing a powerful AI agent a directive any of us might give: *wipe out all cancer. *So the AI begins eliminating every organism capable of developing cancer, which is nearly all life on Earth, *including us. *It didn’t disobey. It didn’t malfunction. Its logic was flawless. It merely misunderstood our true intentions.
We Didn’t Evolve to Understand Our Own Warnings #
We keep warning ourselves with different versions of the same story that we keep ignoring. Why? The answer lives in our roots.
Our brains were built to spot snakes in the grass and anger in a face across the fire, not the challenges of a rapidly changing technological world. Scientists call this evolutionary mismatch, and it sits at the root of many of our cognitive biases. With the progress we evolved to pursue, we’ve created an alien world we didn’t evolve to inhabit.
Exponential change is where the mismatch strikes hardest. Show us a curve that doubles, and we badly underestimate where it lands, an error psychologists call exponential growth bias. We also avoid looking at what we don’t wish to see, so when a threat feels distant and overwhelming, we stop checking on it—the ostrich effect. And our hunter-gatherer ancestors worried about surviving the day, not the century, so we discount the future steeply. Temporal discounting explains why we struggle to save for retirement. It also explains why we fail to register that our technologies are improving exponentially.
I call all of this our evolutionary blindness.
Why We Believe It Anyway #
Human beings do not tolerate open-ended uncertainty about existential threats. We reach for a firm answer and lock onto it, a drive researchers call the need for cognitive closure. The mechanism has a name: *seizing and freezing. *We seize whatever comforting story is available and freeze there. We tell ourselves, "Someone is in charge. It’ll work out, because it always has."
That’s our control delusion. It's not a failure of our intelligence. It reflects a mismatch between the simple, but brutal, world we evolved to live in and the sci-fi world we've created for ourselves.
Seeing Our Delusion Clearly #
We are racing to build something smarter than we are, woven into the digital infrastructure our lives depend on, and telling ourselves we’ll keep it under permanent human control. There’s a reason the animals are inside the cages at the zoo and we humans are on the outside.
The certainty that we can control all of this, indefinitely, is itself the delusion.
The Mirror We Keep Missing #
Losing control of the machines is a legitimate worry, but it may not be the biggest one. The other danger doesn’t require AI to want anything at all. It only requires people. Someone defeated the safeguards on Anthropic’s most powerful model (Fable) within days of release. Thousands of open-weight models circulate with no safeguards to defeat. Governments are building autonomous weapons on purpose. And when governments pulled dangerous models offline, they weren’t afraid of machines going rogue. They were afraid of us.
Artificial IntelligenceEssential Reads Most of us can’t picture what a bad actor will invent, because we don’t spend our days thinking about how to hurt people. We assume others are wired more or less like us, an error known as the false consensus effect. Anaïs Nin captured it: “We don’t see things as they are. We see them as we are.” Call our inability to imagine what bad actors will do with AI a goodness blind spot.
Philosopher Shannon Vallor argues that these systems are mirrors, trained on oceans of our own data, reflecting back the same errors and failures of wisdom we keep hoping to escape (The AI Mirror, 2024). Our technologies act as both mirror and lens, reflecting and magnifying the best and worst in us.
AI can even magnify our evolutionary blindness. But knowing we have blind spots is how we begin to see. Ask AI the right questions, and we might use it as corrective lenses for the very blindness it exploits.
This Is Our Time to Be Unsinkable #
Thoreau saw the shape of this in 1854, writing that our inventions are “improved means to an unimproved end.” Our technology improves exponentially. We don’t. That gap is the whole problem, and it isn’t a problem about machines.
Every retelling of the *Titanic *is haunted by *“if only.” *What makes this moment strange is that we are standing inside our own “if only” before the fact instead of after. There was a narrow corridor through that ice field. A corridor isn’t a straight line, and it doesn’t stay put, so it must be navigated slowly, with every lookout posted. And the icebergs hardest to steer around are the ones we make out of each other.
Now that we know there are icebergs ahead that we can’t see, we can work together to avoid them. Learning from the cautionary tale of the *Titanic *is how we ensure those lives were not lost in vain.
References
Anthropic. (2026). Claude Mythos Preview system card.
Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press.
Frederick, S., Loewenstein, G., & O'Donoghue, T. (2002). Time discounting and time preference: A critical review. Journal of Economic Literature, 40(2), 351–401.
Hugging Face. (2026). Anatomy of a frontier lab agent intrusion: A technical timeline.
Karlsson, N., Loewenstein, G., & Seppi, D. (2009). The ostrich effect: Selective attention to information. Journal of Risk and Uncertainty, 38(2), 95–115.
Kruglanski, A. W., & Webster, D. M. (1996). Motivated closing of the mind: "Seizing" and "freezing." Psychological Review, 103(2), 263–283.
Li, N. P., van Vugt, M., & Colarelli, S. M. (2018). The evolutionary mismatch hypothesis: Implications for psychological science. Current Directions in Psychological Science, 27(1), 38–44.
Ross, L., Greene, D., & House, P. (1977). The "false consensus effect": An egocentric bias in social perception and attribution processes. Journal of Experimental Social Psychology, 13(3), 279–301.
Vallor, S. (2024). The AI Mirror: How to reclaim our humanity in an age of machine thinking. Oxford University Press.
Wagenaar, W. A., & Sagaria, S. D. (1975). Misperception of exponential growth. Perception & Psychophysics, 18(6), 416–422.