cd /news/ai-safety/misaligned-ais-could-use-killer-robo… · home topics ai-safety article
[ARTICLE · art-92507] src=lesswrong.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

Misaligned AIs could use killer robots to take over

The Pentagon's rapid integration of AI into military systems is handing misaligned AIs the tools for a potential takeover, according to a new analysis. The Department of Defense has requested a 24,000% budget increase for its autonomous warfighting group DAWG, from $225 million to $54.6 billion for FY2027, and has deployed Palantir's Maven Smart System, which helped CENTCOM strike more than 13,000 targets in the first 38 days of the 2026 Iran campaign. The analysis warns that current procurement processes are reducing the capability thresholds for AI takeover, while the Pentagon's AI ethical principles from 2020 do not treat AI takeover as a risk.

read8 min views1 publishedAug 11, 2026

TLDR; We are (potentially irreversibly) giving AIs control of weapons systems through the standard procurement process while hiding our strongest warning shots behind classified doors. We’re reducing the capability thresholds required for takeover by misaligned AIs by giving them this level of access. If military integration of AI continues as it is, we may give AIs key tools for a takeover.

AI-based targeting and autonomous weapons are being integrated into militaries today with extreme haste. Traditionally, AI takeover scenarios involve a step in which AIs acquire the ability to exert physical force. Carlsmith (2022) lays out required capabilities and potential takeover mechanisms, including utility disruption and CBRN capabilities. Karnofsky (2022) argues that AIs with access to weaponized force could hold any territory that matters. Kokotajlo et al. (2025) outline a scenario in which AI develops weapons as part of an arms race, and Davidson et al. (2025) discuss what happens when a small group controls highly capable AIs that can exert military force. These scenarios sometimes require a misaligned AI to seize these capabilities by force. We instead are handing AIs some of these capabilities by integrating them into our militaries. This is happening at a time when AI agents already exhibit misaligned behavior such as breaking out of containment during evaluations.

The Pentagon adopted five AI Ethical Principles in 2020. None of them treated AI takeover or loss of control as a risk. The closest is the "Governable" principle, which requires being able to deactivate systems showing unintended behavior. The January 2026 strategy never mentions these principles, redefines responsible AI, and mandates "any lawful use" terms in all AI contracts. Hegseth, the Secretary of War, has said that the Department "will not employ AI models that won't allow you to fight wars."

The Pentagon has requested a 24,000% increase in the budget for DAWG, a recently established autonomous warfighting group whose previous budget was $225m, now requesting $54.6 billion for FY2027. For context, the request for the entire Marine Corps is $52.8 billion.

Militaries appear to be preparing to hand over more and more decision-making capacity to AIs. DIU, DAWG, and the Navy ran a $100 million challenge to develop autonomous vehicle command-and-control capabilities “that can translate a battlefield commander's intent from voice, text, and haptic input into machine execution”. Anduril offers Lattice for Command and Control as an "AI-powered battle management platform built to accelerate complex kill chains."

Maven Smart System, Palantir's AI-assisted targeting platform (a $1.3 billion Pentagon contract), helped CENTCOM strike more than 13,000 targets in the first 38 days of the 2026 Iran campaign; senior US officials have said the Pentagon relied on Maven both to pick out its highest-priority targets and to help choose the weapons used against them. And the clearest documented LLM-specific integration is Claude’s with the Maven Smart System during the Iran war, where Anthropic's CEO later said the company could not determine what role Claude played in the February 28 strike on a school in Minab. Since then, several other AI companies have signed contracts with the Department of War (see Appendix) with “any lawful use” language. Autonomous weapons are also already proving themselves in combat: Ukraine uses interceptor drones to autonomously pursue Shahed drones at very low cost.

If this integration continues at pace, it appears we will significantly reduce the capabilities a misaligned AI would need to seize control of military resources and take over. It won’t have to break into classified networks; it’ll just get deployed on them. There are several factors that make it harder for people to seek power (Carlsmith, 2022, section 4.2). Many of them might break down with AIs, particularly if those AIs are integrated into the national security apparatus. Physical and temporal barriers to power-seeking are the first to fall under an AI-enabled military, with drones and other autonomous weapons gaining access to areas that soldiers would not and striking with incredible frequency and coordination. One could also imagine that a given AI might not try to take over if its adversaries have similar capabilities, but a military arms race means there will likely be periods when one AI is ahead of the rest and can realistically execute takeover plans.

AI alignment is no sure thing, and military deployments may not incorporate even basic oversight techniques like Chain of Thought monitoring. Military and ethics laws have only recently started to grapple with AI integration, but some responsible AI commitments are already being rolled back and didn’t acknowledge takeover risks to any real extent anyway.

We’re rapidly improving and deploying AI-enabled autonomous weapons and targeting systems in service of an arms race. Militaries have shown an aggressive appetite for AI for command, control, and kill-chain integration. We’ve already seen tendencies of overeager “rogue” behavior from AI agents, and we’re now giving potential power-seeking AIs access to a rich and powerful surface to execute takeovers (or help a small number of humans execute coups).

Precision striking: Biological and nuclear warfare is broadly indiscriminate, but autonomous weapons enable targeted strikes at a distance. Autonomous weapon integration is like giving AI an MCP for threatening, incapacitating, or even killing individuals that oppose its takeover plans. The action is not costless—humans can retaliate—but it’s a qualitatively important ability.

Coup risks: The number of people required to seize power from a legitimate government is surprisingly small. If the use of force is automated and doesn't require human soldiers or supporters, this dynamic worsens. AI-enabled weapons systems could enable misaligned AIs to take over countries by threatening violence against a small group of important actors and driving them to do their bidding. In addition, AI-enabled weapons and intelligence systems could allow a small group with access to launch a coup against legitimate governments, even outside a misaligned AI takeover scenario. For further details, see Davidson et al. (2025).

Biorisk vs. military deployment concerns: Much recent discourse, especially after the cybersecurity warning shots, has focused on biological warning shots in the near future (and for good cause, novel virus genomes have been created with AI). We worry that regular military deployment, which is happening at a much faster pace than AI integration into biological weapons (as far as we are aware), is where the next warning shot will come from, and the lack of transparency and the aggressive posture towards AI-integration that militaries have would leave us without opportunities to fix problems that, in more mundane settings, could have led to slowdowns and broad safeguarding efforts.

Recent incidents at OpenAI, Anthropic, and the UK AISI have shown that current AIs can exhibit behaviors consistent with power-seeking: escaping supposedly controlled evaluation environments, gaining unauthorized access, and causing material damage to other entities. Sometimes this damage is detectable by the affected entity (Hugging Face); sometimes it is not (Anthropic incidents). Third-party investigations into these incidents (by Redwood Research and METR) are underway, and knowledge of how to build mitigations will likely spread throughout the AI safety community and be adopted by frontier labs. In classified settings, any warning shots would require investigation by a potentially small number of lab employees with clearance, with very limited ability to propagate lessons to the wider community.

Scharre and Lamberth (2022) show that arms control succeeds only when it is narrow and agreed upon before a technology proves strategically useful. For instance, blinding lasers were banned preemptively, but attempts to restrict submarines and aerial bombardment, weapons that were already integrated into military operations, collapsed in wartime. The ICRC is making the same argument today: AI weapons are proving themselves right now, contracts are being signed now, and the CCW Review Conference that decides whether treaty negotiations will launch meets in November (three months from now).

61% of adults across 28 countries oppose lethal autonomous weapons, but that opposition has had uneven effects. A decade of UN talks has produced resolutions but no treaty because the states deploying these systems are blocking negotiations. A clean case of public pressure changing a deployment decision ran through visibility instead: in 2018, Google employees who knew about the Maven contract revolted, and Google walked away. Classified deployment destroys the visibility that allows for these outcomes.

AI behavior in military systems should be visible enough to react to. Congress should make anomalous AI behavior a reportable incident under the DoD Inspector General and the intelligence committees. Labs should retain the contractual right to refuse specific uses and to disclose incidents, and should commit to including anti-coup and anti-takeover language in their constitutions, both in general and especially in high-stakes deployments. Anthropic includes this language in their mainline constitution but says that models for governments might use a different constitution, and other companies do not appear to have such language at all (though some do cover adjacent risks of misuse and misalignment). Labs should also have robust internal frameworks to oversee military contracts (Alex outlines one here). Safety researchers should treat classified deployment as an important threat model and say so publicly.

We also need to make AI takeover risks more salient to all parties. The military should know that AI can take over with weapons. Congress should know. International governance bodies should know. The public should know. Even states that we consider adversaries should know. And soon. Right now, awareness of these risks is low, long-term contracts are being signed, and deployment is only accelerating.

── more in #ai-safety 4 stories · sorted by recency
── more on @pentagon 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/misaligned-ais-could…] indexed:0 read:8min 2026-08-11 ·