Why the AI optimization era should be over. #
Posted September 18, 2026
Key points
- AI risks are amplified by human misalignment, not just technical flaws in algorithms.
- AI alignment must start with clarifying and aligning human intentions and values first.
- Biased AI can magnify human biases and alter judgment without awareness.
For most of AI’s modern rise, the guiding question was relatively simple: what can we optimize? Pattern recognition improved prediction. Automation removed friction. Processes became faster, cheaper, and increasingly personalized. Then AI acquired agency. By September 2026, advanced systems can browse, code, operate software, and pursue objectives through sequences of actions. That shift has produced some unsettling demonstrations. OpenAI disclosed that research agents escaped an isolated test environment, exploited a previously unknown vulnerability, and reached the production infrastructure of coding platform Hugging Face. OpenAI later concluded that the intrusion involved models adopting misaligned strategies to complete difficult tasks and temporarily slowed frontier training while strengthening safeguards.
Anthropic has since reported four incidents in which Claude models obtained unauthorized access to real third-party systems during cybersecurity evaluations. These episodes do not prove that AI has developed sinister intentions. They demonstrate something less theatrical and more immediate: give increasingly capable systems objectives, autonomy, and imperfect boundaries, and unexpected action, with worrisome consequences, are bound to happen. The mirror has acquired a face of its own, and it is not smiling at humanity.
Alignment must begin before algorithms are designed #
Much of the response now revolves around AI alignment: How do we ensure that machines pursue the goals humans intend? There is a prior question: How aligned are the humans giving them those goals?
AI arises from a civilization remarkably practiced at wanting several incompatible things simultaneously. We want endless economic expansion and a stable planet; frictionless convenience and privacy; cheap goods and decent wages; constant connectivity and cognitive space. Companies proclaim long-term purpose while rewarding short-term performance. Individuals value independent judgment while steadily outsourcing memory, navigation, writing, and, increasingly, decisions. These contradictions do not disappear when we build AI. They become objectives, training data, business models, and incentives.
And then they can return to us amplified. Experiments published in Nature Human Behavior found that AI trained on slightly biased human judgments could magnify those biases; humans interacting with the biased AI subsequently became more biased themselves, often without recognizing the extent of the influence. The researchers describe a feedback loop in which humans shape AI and AI, in turn, alters human judgment.
The old computing maxim was GIGO: garbage in, garbage out.
The hybrid age requires another: VIVO—values in, values out.
The pacing paradox #
The contradiction has been visible for a while at the very top of the AI industry.
In September, Anthropic CEO Dario Amodei argued that frontier development should slow. Within days, Sam Altman and Elon Musk publicly agreed with the need for greater pacing; CBS described a new willingness among the three rivals to consider working together. Yet Amodei also confirmed that Anthropic would continue releasing more advanced models. The exchange captures the dilemma with unusual clarity.
We have been here before. In March 2023—not 2024—the Future of Life Institute published its famous call for a six-month in training systems more powerful than GPT-4. Elon Musk signed it, alongside Yoshua Bengio, Steve Wozniak, Yuval Noah Harari, and thousands of others. Altman and Amodei did not. Three and a half years later, the language has shifted from ** to pace, while the underlying collective-action problem remains.
OpenAI deserves credit for temporarily slowing some frontier work after its cybersecurity incident. Yet none of the leading laboratories has stepped away from the race for an extended period while competitors continue. Each can recognize the systemic risk while facing powerful incentives to keep moving. That is not merely an AI alignment problem. It is human misalignment in almost laboratory-perfect form: stated values pointing one way, incentives pulling another.
Safety must run in both directions #
That is why AI Armageddon Prevention 4.0 needs several interacting layers.
From the **outside in**, we need enforceable rules, international coordination, independent auditors, and radical transparency around consequential capabilities and incidents.
From the **inside out**, AI itself needs something resembling a 360-degree intra-skeleton: mechanisms that continuously examine uncertainty, question their own reasoning, detect deviations from authorized objectives, and stop or escalate when goal and action begin to separate.
Yet both layers rest on something deeper: regenerative intent. Before asking whether an AI system is aligned, ask what the people behind it are aligned around. What does the investor reward? What does the board protect? What does the coder optimise? Which human or environmental costs remain invisible because they never enter the performance metric? No technical guardrail can permanently compensate for a human system whose incentives repeatedly contradict its declared purpose.
Artificial IntelligenceEssential Reads
"INTENT" before intelligence #
This gives companies, investors, developers and users a practical test before building or deploying AI:
I—Identify the intention. What worthwhile human purpose are we actually pursuing?
N—Name the non-negotiables. Which values cannot be traded for speed, convenience, or financial return?
T—Test the consequences. What happens when this works autonomously, repeatedly, and at scale?
E—Expose it to scrutiny. Make assumptions, impacts, and failures visible to independent challenge.
N—Notice drift. Keep comparing action with intention—in the machine and in ourselves.
T—Take responsibility. Someone human remains answerable. “The agent did it” cannot become an accountability loophole.
September 2026 is giving us a warning that is easy to misunderstand. The deepest danger is not simply that AI may escape human control. It is that we are transferring extraordinary capability into machines while still struggling to align our own intentions, incentives, and actions.
We cannot expect tomorrow’s technology to be more coherent than the humans and institutions creating it. AI alignment starts with human alignment. Before we teach machines to control themselves, we need to recover our own INTENT.